For AI agents: a documentation index is available at https://www.mongodb.com/docs/llms.txt — markdown versions of all pages are available by appending .md to any URL path.
Docs Menu

Perform Parent Document Retrieval with MongoDB and LangChain

You can integrate MongoDB Vector Search with LangChain to perform parent document retrieval. In this tutorial, you complete the following steps:

  1. Set up the environment.

  2. Prepare the data.

  3. Instantiate the parent document retriever.

  4. Create the MongoDB Vector Search index.

  5. Use the retriever in a RAG pipeline.

Parent document retrieval is a retrieval technique that involves chunking large documents into smaller sub-documents. In this technique, you query the smaller chunks before returning the full parent document to the LLM. This can improve the responses of your RAG agents and applications by allowing for more granular searches on smaller chunks while giving LLMs the full context of the parent document.

Parent document retrieval with MongoDB allows you to store both parent and child documents in a single collection, which supports efficient retrieval by only having to compute and index the child documents' embeddings.

To complete this tutorial, you must have the following:

  • One of the following MongoDB cluster types:

    • An Atlas cluster running MongoDB version 6.0.11, 7.0.2, or later. Ensure that your IP address is included in your Atlas project's access list.

    • A local Atlas deployment created using Python and Docker. Install atlas-local-lib-py (pip install atlas-local-lib-py) to programmatically create and manage local deployments. To learn more, see the atlas-local-lib-py repository.

    • A MongoDB Community cluster with Search and Vector Search installed.

  • A Voyage AI API key. To create an API key, see Manage Voyage AI Model API Keys.

  • An OpenAI API Key. You must have an OpenAI account with credits available for API requests. To learn more about registering an OpenAI account, see the OpenAI API website.

  • An environment to run interactive Python notebooks such as Colab.

Set up the environment for this tutorial. Create an interactive Python notebook by saving a file with the .ipynb extension. This notebook allows you to run Python code snippets individually, and you'll use it to run the code in this tutorial.

To set up your notebook environment:

1

Run the following command:

pip install --quiet --upgrade langchain langchain-community langchain-core langchain-mongodb langchain-voyageai langchain-openai pymongo pypdf
2

Run the following code, replacing the placeholders with the following values:

  • Your Voyage AI and OpenAI API Key.

  • Your MongoDB cluster's connection string.

import os
os.environ["VOYAGE_API_KEY"] = "<voyage-api-key>"
os.environ["OPENAI_API_KEY"] = "<openai-api-key>"
MONGODB_URI = "<connection-string>"

Note

Replace <connection-string> with the connection string for your Atlas cluster or local Atlas deployment.

Your connection string should use the following format:

mongodb+srv://<db_username>:<db_password>@<clusterName>.<hostname>.mongodb.net

To learn more, see Connect to a Cluster via Client Libraries.

Your connection string should use the following format:

mongodb://localhost:<port-number>/?directConnection=true

To learn more, see Connection Strings.

Paste and run the following code in your notebook to load and chunk a sample PDF that contains a recent MongoDB earnings report.

This code uses a text splitter to chunk the PDF data into smaller parent documents. It specifies the chunk size (number of characters) and chunk overlap (number of overlapping characters between consecutive chunks) for each document.

from langchain_text_splitters import RecursiveCharacterTextSplitter
from langchain_community.document_loaders import PyPDFLoader
# Load the PDF
loader = PyPDFLoader("https://investors.mongodb.com/node/12881/pdf")
data = loader.load()
# Chunk into parent documents
parent_splitter = RecursiveCharacterTextSplitter(chunk_size=2000, chunk_overlap=20)
docs = parent_splitter.split_documents(data)
# Print a document
docs[0]

In this section, you instantiate the parent document retriever and use it to ingest data into MongoDB.

MongoDBAtlasParentDocumentRetriever chunks parent documents into smaller child documents, embeds the child documents, and then ingests both parent and child documents into the same collection in MongoDB. Under the hood, this retriever creates the following:

  • An instance of MongoDBAtlasVectorSearch, a vector store that handles vector search queries to the child documents.

  • An instance of MongoDBDocStore, a document store that handles storing and retrieving the parent documents.

1

The fastest way to configure MongoDBAtlasParentDocumentRetriever is to use the from_connection_string method. This code specifies the following parameters:

  • connection_string: Your Atlas connection string to connect to your cluster.

  • child_splitter: The text splitter to use to split the parent documents into smaller, child documents.

  • embedding_model: The embedding model to use to embed the child documents.

  • database_name and collection_name: The database and collection name for which to ingest the documents.

  • The following optional parameters to configure the MongoDBAtlasVectorSearch vector store:

    • text_key: The field in the documents that contains the text to embed.

    • relevance_score: The relevance score to use for the vector search query.

    • search_kwargs: How many child documents to retrieve in the initial search.

from langchain_mongodb.retrievers import MongoDBAtlasParentDocumentRetriever
from langchain_voyageai import VoyageAIEmbeddings
# Define the embedding model to use
embedding_model = VoyageAIEmbeddings(model="voyage-3-large")
# Define the chunking method for the child documents
child_splitter = RecursiveCharacterTextSplitter(chunk_size=200, chunk_overlap=20)
# Specify the database and collection name
database_name = "langchain_db"
collection_name = "parent_document"
# Create the parent document retriever
parent_doc_retriever = MongoDBAtlasParentDocumentRetriever.from_connection_string(
connection_string = MONGODB_URI,
child_splitter = child_splitter,
embedding_model = embedding_model,
database_name = database_name,
collection_name = collection_name,
text_key = "page_content",
relevance_score_fn = "dotProduct",
search_kwargs = { "k": 10 },
)
2

Then, run the following code to ingest the documents into Atlas by using retriever's add_documents method. It takes the parent documents as an input and ingests both parent and child documents based on how you configured the retriever.

parent_doc_retriever.add_documents(docs)
3

After running the sample code, you can view the documents in the Atlas UI by navigating to the langchain_db.parent_document collection in your cluster.

Both parent and child documents have a page_content field that contains the chunked text. The child documents also have an additional embedding field that contains the vector embeddings of the chunked text, and a doc_id field that corresponds to the _id of the parent document.

You can run the following queries in the Atlas UI, replacing the <id> placeholder with a valid document ID:

  • To see child documents that share the same parent document ID:

    { doc_id: "<id>" }
  • To see the parent document of those child documents:

    { _id: "<id>" }

To enable vector search queries on the langchain_db.parent_document collection, you must create a MongoDB Vector Search index. You can use either the LangChain helper method or the PyMongo driver method. Run the following code in your notebook for your preferred method:

# Get the vector store instance from the retriever
vector_store = parent_doc_retriever.vectorstore
# Use helper method to create the vector search index
vector_store.create_vector_search_index(
dimensions = 1024 # The number of dimensions to index
)
from pymongo import MongoClient
from pymongo.operations import SearchIndexModel
# Connect to your cluster
client = MongoClient(MONGODB_URI)
collection = client[database_name][collection_name]
# Create your vector search index model, then create the index
vector_index_model = SearchIndexModel(
definition={
"fields": [
{
"type": "vector",
"path": "embedding",
"numDimensions": 1024,
"similarity": "dotProduct"
}
]
},
name="vector_index",
type="vectorSearch"
)
collection.create_search_index(model=vector_index_model)

The index should take about one minute to build. While it builds, the index is in an initial sync state. When it finishes building, you can start querying the data in your collection.

Once MongoDB builds your index, you can run vector search queries on your data and use the retriever in your RAG pipeline. Paste and run the following code in your notebook to implement a sample RAG pipeline that performs parent document retrieval:

1

To see the most relevant documents for a given query, paste and run the following code to perform a sample vector search query on the collection. The retriever searches for relevant child documents that are semantically similar to the string AI technology, and then returns the corresponding parent documents of the child documents.

parent_doc_retriever.invoke("AI technology")

To learn more about vector search query examples with LangChain, see Run Vector Search Queries.

2

To create and run a RAG pipeline with the parent document retriever, paste and run the following code. This code does the following:

  • Defines a LangChain prompt template to instruct the LLM to use the retrieved parent documents as context for your query. LangChain passes these documents to the {context} input variable and your query to the {query} variable.

  • Constructs a chain that specifies the following:

    • The parent document retriever you configured to retrieve relevant parent documents.

    • The prompt template that you defined.

    • An LLM from OpenAI to generate a context-aware response. By default, this is the gpt-3.5-turbo model.

  • Prompts the chain with a sample query and returns the response. The generated response might vary.

from langchain_core.output_parsers import StrOutputParser
from langchain_core.prompts import PromptTemplate
from langchain_core.runnables import RunnablePassthrough
from langchain_openai import ChatOpenAI
# Define a prompt template
template = """
Use the following pieces of context to answer the question at the end.
{context}
Question: {query}?
"""
prompt = PromptTemplate.from_template(template)
model = ChatOpenAI()
# Construct a chain to answer questions on your data
chain = (
{"context": parent_doc_retriever, "query": RunnablePassthrough()}
| prompt
| model
| StrOutputParser()
)
# Prompt the chain
query = "In a list, what are MongoDB's latest AI announcements?"
answer = chain.invoke(query)
print(answer)
1. MongoDB obtained the AWS Modernization Competency designation.
2. MongoDB launched a MongoDB University course focused on building AI applications with MongoDB and AWS.
3. MongoDB announced new technology integrations for AI, data analytics, and automating database deployments across various environments.
4. MongoDB launched the MongoDB AI Applications Program (MAAP) to help companies harness the power of data and future AI technologies.
5. Capgemini, Confluent, IBM, Unstructured, and QuantumBlack joined the MAAP ecosystem to offer customers additional integration and solution options.

Follow along with this video about parent document retrieval with LangChain and MongoDB.

Duration: 27 Minutes