MongoDB をLgChuin と統合して、生成系AIと RAG アプリケーションを構築できます。このページでは、LgChuin MongoDB Python統合とアプリケーションで使用できるさまざまなコンポーネントの概要について説明します。
はじめるインストールとセットアップ
LgChuin でMongoDB ベクトル検索を使用するには、まず langchain-mongodbパッケージをインストールする必要があります。
pip install langchain-mongodb
ベクトル ストア
MongoDBAtlasVectorSearch は、 MongoDBのコレクションからベクトル埋め込みを保存および検索できるベクトルストアです。このコンポーネントを使用してデータの埋め込みを保存し、 MongoDB ベクトル検索を使用して埋め込みを検索できます。
このコンポーネントにはMongoDB ベクトル検索インデックスが必要です。
使用法
Atlas は 2 つの埋め込みモードをサポートしています。
手動埋め込み: 指定した埋め込みモデルを使用して、クライアント側で埋め込みベクトルを生成します。
自動埋め込み: MongoDB は、手動で生成する必要なく、サーバー側にテキストを埋め込みます。詳細については、「 自動埋め込み 」を参照してください。
Parameter | 必要性 | 説明 |
|---|---|---|
| 必須 | MongoDBクラスターの接続文字列を指定します。詳細については、クライアント ライブラリを使用したクラスターへの接続 または 接続文字列 を参照してください。 |
| 必須 | ベクトル埋め込みを保存するための MongoDB 名前空間を指定してください。 たとえば、 |
| 必須 | |
| 任意 | MongoDB ベクトル検索インデックスの名前。デフォルトは |
| 任意 | ドキュメントのテキスト コンテンツを含むフィールド名。デフォルトは |
| 任意 | 埋め込みベクトルを保存するフィールド名。デフォルトは |
| 任意 | 使用する類似度関数。使用可能な値は |
| 任意 | ベクトル次元の数。この値を設定しており、コレクションにベクトル検索インデックスがない場合は、 MongoDB がインデックスを作成します。 |
| 任意 | ベクトル インデックスがない場合に、自動的に作成するかどうかを決定するフラグ。デフォルトは |
| 任意 | 自動生成されたベクトル検索インデックスの準備完了を待つ際のタイムアウト時間(秒)。 |
| 任意 | ベクトル検索インデックスを構成するための追加オプションの辞書。 |
| 任意 | LangChain 固有のパラメーターなど、ベクトル ストアに渡す追加のパラメーター。 |
Retrievers
LangChain 検索 は、ベクトルストアから関連するドキュメントを取得するために使用するコンポーネントです。LgChuin に組み込まれている検索ドライバーまたは次のMongoDB検索ドライバーを使用して、 MongoDBからデータをクエリして検索できます。
ベクトル検索リージョン
MongoDB をベクトルストアとしてインスタンス化したら、ベクトルストアのインスタンスを検索用に使用して、MongoDB ベクトル検索を使用してデータをクエリできます。
使用法
from langchain_mongodb.vectorstores import MongoDBAtlasVectorSearch from langchain_voyageai import VoyageAIEmbeddings # Instantiate the vector store vector_store = MongoDBAtlasVectorSearch.from_connection_string( connection_string="<connection-string>", # MongoDB cluster URI namespace="<database-name>.<collection-name>", # Database and collection name embedding=VoyageAIEmbeddings(model="voyage-3-large"), # Embedding model to use index_name="vector_index", # Name of the vector search index ) # Use the vector store as a retriever retriever = vector_store.as_retriever() # Define your query query = "some search query" # Print results documents = retriever.invoke(query) for doc in documents: print(doc)
全文検索システム
MongoDBAtlasFullTextSearchRetriever は、 MongoDB Search を使用して全文検索を実行する検索ドライバーです。具体的には、 Lucene の標準 IBM25 アルゴリズムを使用します。
この検索インデックスにはMongoDB Search インデックスが必要です。
使用法
from langchain_mongodb.retrievers.full_text_search import ( MongoDBAtlasFullTextSearchRetriever, ) from pymongo import MongoClient # Connect to your MongoDB cluster client = MongoClient("<connection-string>") collection = client["<database-name>"]["<collection-name>"] # Initialize the retriever retriever = MongoDBAtlasFullTextSearchRetriever( collection=collection, # MongoDB Collection in Atlas search_field="<field-name>", # Name of the field to search search_index_name="<index-name>", # Name of the search index ) # Define your query query = "some search query" # Print results documents = retriever.invoke(query) for doc in documents: print(doc)
注意
ハイブリッド検索システム
MongoDBAtlasHybridSearchRetriever は、 レプリカ ランク統合(RRF)アルゴリズムを使用して、ベクトル検索と全文検索の結果を組み合わせた検索結果です。「ハイブリッド検索の実行方法」を参照してください。
この検索インデックスには、既存のベクトルストア、MongoDB ベクトル検索インデックス、およびMongoDB Search インデックスが必要です。
使用法
from langchain_mongodb.retrievers.hybrid_search import ( MongoDBAtlasHybridSearchRetriever, ) from langchain_mongodb.vectorstores import MongoDBAtlasVectorSearch from langchain_voyageai import VoyageAIEmbeddings # Instantiate the vector store vector_store = MongoDBAtlasVectorSearch.from_connection_string( connection_string="<connection-string>", # MongoDB cluster URI namespace="<database-name>.<collection-name>", # Database and collection name embedding=VoyageAIEmbeddings(model="voyage-3-large"), # Embedding model to use index_name="vector_index", # Name of the vector search index ) # Initialize the retriever retriever = MongoDBAtlasHybridSearchRetriever( vectorstore=vector_store, # Vector store instance search_index_name="<index-name>", # Name of the MongoDB Search index top_k=5, # Number of documents to return fulltext_penalty=60.0, # Penalty for full-text search vector_penalty=60.0, # Penalty for vector search ) # Define your query query = "some search query" # Print results documents = retriever.invoke(query) for doc in documents: print(doc)
Parent Document Retriever
MongoDBAtlasParentDocumentRetriever は、最初に小さなチャンクをクエリし、その後、大きな親ドキュメントを LLM に返す検索システムです。このタイプの検索システムは、親ドキュメント検索と呼ばれます。親ドキュメント検索により、より小さなチャンクでのより詳細な検索が可能になり、同時に LLM に親ドキュメントの完全なコンテキストが提供されるため、RAG エージェントとアプリケーションの応答が向上します。
このレトリーバーは、親ドキュメントと子ドキュメントの両方を 1 つの MongoDB コレクションに保存できるため、子ドキュメントの埋め込みを計算してインデックスを作成するだけで効率的な検索が可能になります。
内部的には、このレトリーバーは以下を作成します。
子ドキュメントに対するベクトル検索クエリを取り扱う MongoDBAtlasVectorSearch のインスタンス。
親ドキュメントの保存と検索を取り扱う MongoDBDocStore のインスタンス。
使用法
text_key を page_content に設定して、ベクトルストアと親ドキュメントストアがドキュメントテキストに同じフィールド名を使用するようにします。このパラメータを指定しない場合、検索ドライバーは親ドキュメントをあるフィールドに書き込み、別のフィールドから読み取り、クエリは KeyError: 'text' で失敗します。
from langchain_mongodb.retrievers import MongoDBAtlasParentDocumentRetriever from langchain_text_splitters import RecursiveCharacterTextSplitter from langchain_voyageai import VoyageAIEmbeddings retriever = MongoDBAtlasParentDocumentRetriever.from_connection_string( connection_string="<connection-string>", # MongoDB cluster URI embedding_model=VoyageAIEmbeddings( # Embedding model to use model="voyage-3-large" ), child_splitter=RecursiveCharacterTextSplitter(), # Text splitter to use database_name="<database-name>", # Database to store the collection collection_name="<collection-name>", # Collection to store the collection text_key="page_content", # Match the key the parent document store uses # Additional vector store or parent class arguments... ) # Define your query query = "some search query" # Print results documents = retriever.invoke(query) for doc in documents: print(doc)
自己クエリ検索プロセス
MongoDBAtlasSelfQueryRetriever は、それ自体をクエリするレプリカです。検索クエリークエリーは LM を使用して処理され、可能なメタデータフィルターを識別し、フィルター付きで構造化ベクトル検索クエリーを作成し、そのクエリを実行して最も関連性の高いドキュメントを検索します。
例、「2010 以降の評価が 8 を超えるアクション映画は何ですか」のようなクエリでは、取得者は genre、year、rating フィールドのフィルターを識別し、それらを使用できますクエリに一致するドキュメントを検索するためにフィルタリングします。
この検索インデックスには、既存のベクトルストアとMongoDB ベクトル検索インデックスが必要です。
使用法
from langchain_mongodb.retrievers import MongoDBAtlasSelfQueryRetriever from langchain_mongodb import MongoDBAtlasVectorSearch from langchain_classic.chains.query_constructor.schema import AttributeInfo from langchain_voyageai import VoyageAIEmbeddings from langchain_openai import ChatOpenAI llm = ChatOpenAI(model="gpt-4o", temperature=0) vector_store = MongoDBAtlasVectorSearch.from_connection_string( connection_string="<connection-string>", namespace="langchain_db.movies", embedding=VoyageAIEmbeddings(model="voyage-3-large"), index_name="vector_index", ) # Given an existing vector store with movies data, define metadata describing the data metadata_field_info = [ AttributeInfo( name="genre", description="The genre of the movie. One of ['science fiction', 'comedy', 'drama', 'thriller', 'romance', 'animated']", type="string", ), AttributeInfo( name="year", description="The year the movie was released", type="integer", ), AttributeInfo( name="rating", description="A 1-10 rating for the movie", type="float" ), ] # Create the retriever from the VectorStore, an LLM and info about the documents retriever = MongoDBAtlasSelfQueryRetriever.from_llm( llm=llm, vectorstore=vector_store, metadata_field_info=metadata_field_info, document_contents="Descriptions of movies", enable_limit=True, ) # This example results in the following composite filter sent to $vectorSearch: # {'filter': {'$and': [{'year': {'$lt': 1960}}, {'rating': {'$gt': 8}}]}} documents = retriever.invoke("Movies made before 1960 that are rated higher than 8") print(documents)
GraphRAG
GraphRAG は、ベクトル埋め込みとしてではなく、エンティティとその関係の知識グラフとしてデータを構造化する従来の RG の代替アプローチです。ベクトルベースの RG はクエリにセマンティックに類似するドキュメントを検索しますが、GraphRAG はクエリに接続されたエンティティを検索し、グラフ内の関係を走査して関連情報を検索します。
このアプローチは関係がベースとなる質問で特に役立ちます。たとえば、「A 会社と B 会社のつながりは何ですか?」や「X さんのマネージャーは誰ですか?」などの質問です。
MongoDBGraphStore は、LangChain MongoDB 統合内のコンポーネントであり、エンティティ(ノード)とその関係(エッジ)を MongoDB コレクションに保存することで GraphRAG を実装できるようにします。コレクション内の他のドキュメントを参照する関係フィールドのあるドキュメントとして各エンティティを保存し、$graphLookup 集計ステージを使用してクエリを実行します。
使用法
from langchain_mongodb.graphrag import MongoDBGraphStore from langchain_openai import ChatOpenAI from langchain_core.documents import Document # Initialize the graph store graph_store = MongoDBGraphStore( connection_string="<connection-string>", # MongoDB cluster URI database_name="<database-name>", # Database to store the graph collection_name="<collection-name>", # Collection to store the graph entity_extraction_model=ChatOpenAI( # LLM to extract entities model="gpt-4o", temperature=0 ), # Other optional parameters... ) # Add documents to the graph docs = [ Document( page_content=( "MongoDB is a document database. " "Dev Ittycheria is the CEO of MongoDB." ) ), Document(page_content="MongoDB Atlas is the cloud platform offered by MongoDB."), ] graph_store.add_documents(docs) # Query the graph query = "Who is the CEO of MongoDB?" answer = graph_store.chat_response(query) print(answer.content)
LLM キャッシュ
キャッシュ は、類似または反復的なクエリに対する反復的な応答を保存して再計算を回避することで、LLM パフォーマンスを最適化するために使用されます。MongoDB は、LangChain アプリケーションに対して次のキャッシュを提供します。
MongoDB キャッシュ
MongoDBCache を使用すると、 MongoDBコレクションに基本的なキャッシュを保存できます。
使用法
from langchain_mongodb import MongoDBCache from langchain_core.globals import set_llm_cache set_llm_cache( MongoDBCache( connection_string="<connection-string>", # MongoDB cluster URI database_name="langchain_db", # Database to store the cache collection_name="cache", # Collection to store the cache ) )
セマンティックキャッシュ
セマンティックキャッシュは、ユーザー入力とキャッシュされた結果のセマンティックな類似性に基づいて、キャッシュされたプロンプトを検索する、より高度なキャッシュ形式です。
MongoDBAtlasSemanticCache は、 MongoDB ベクトル検索を使用してキャッシュされたプロンプトを検索するセマンティックキャッシュです。このコンポーネントにはMongoDB ベクトル検索インデックスが必要です。
使用法
from langchain_mongodb import MongoDBAtlasSemanticCache from langchain_core.globals import set_llm_cache from langchain_voyageai import VoyageAIEmbeddings set_llm_cache( MongoDBAtlasSemanticCache( embedding=VoyageAIEmbeddings(model="voyage-3-large"), # Embedding model connection_string="<connection-string>", # MongoDB cluster URI database_name="langchain_db", # Database to store the cache collection_name="semantic_cache", # Collection to store the cache ) )
Dep Agents 仮想ファイル システム
LgChuin Desktop は、長時間実行されるマルチステップ タスク用に設計されたエージェントハードウェアです。プランニング、コンテキスト マネジメント、サブエージェントへの作業の委任を取り扱います。シャードは、エージェントのファイルが実際に存在する場所を変更するためにスワップ可能なバックエンドプロトコルをサポートしています。 langchain-mongodb-deepagents-vfsパッケージはそのプロトコルの実装です。Amazon S3はファイルを保持し、埋め込みプロバイダー(AWS Reduce または OpenAI)は埋め込みを計算し、 MongoDB Atlas はチャンクと埋め込みを保持し、MongoFilesystemBackendクラスは各ファイルをルーティングします。ハードウェア メトリクスを操作ようにします。
Use this package when your agent needs to search a large set of existing files in S3. grep runs as a single MongoDB aggregation that combines full-text and vector search results using the Reciprocal Rank Fusion (RRF) algorithm, so it scales without loading every file into your agent to filter them one by one. glob and ls handle filename and directory lookups directly. When an agent calls read, write, edit, upload_files, or download_files, those calls go directly to S3, where your files live. Files added by other tools are automatically picked up and indexed by the backend's watcher.
パッケージをインストールする前に、以下があることを確認してください。
MongoDB Atlas接続文字列
AWS認証情報(
AWS_ACCESS_KEY_ID、AWS_SECRET_ACCESS_KEY、AWS_DEFAULT_REGION)。IAM ポリシーは次の条件を満たす場合s3:GetObject,s3:PutObject,s3:ListBucket, ands3:DeleteObjecton your S3 bucketbedrock:InvokeModelデフォルトのBedlock プロバイダーを使用する場合は、AWS_DEFAULT_REGIONと同じリージョンのamazon.titan-embed-text-v2:0が必要
埋め込みプロバイダーの選択: Redock(デフォルトの、上記のAWS認証情報を使用)または OpenAI (
EMBEDDING_PROVIDER=openaiを設定し、OPENAI_API_KEYを提供)
パッケージをインストールするには、 MongoDB がAWS Bear または OpenAI を使用して検索埋め込みを生成するかを決定し、一致する コマンドを実行します。
pip install "langchain-mongodb-deepagents-vfs[bedrock]"
pip install "langchain-mongodb-deepagents-vfs[openai]"
使用法
S3バケット名と Atlas接続文字列を使用して MongoFilesystemBackend をインスタンス化します。次の例では、 S3 に 2 つのファイルを書き込み、それぞれの検索方法を示しています。
grepは、ファイルの内容を検索しますglobはパターンでファイルパスと一致しますlsディレクトリの内容を一覧表示する
from langchain_mongodb_deepagents_vfs import MongoFilesystemBackend # Instantiate the backend backend = MongoFilesystemBackend( s3_bucket_name="<bucket-name>", # S3 bucket that stores your files mongodb_connection_string="<connection-string>", # MongoDB Atlas connection string ) # Write two files to S3: one .txt, one .md, so glob can demonstrate # filtering by extension backend.write("mongodb_vfs/docs/notes.txt", "Our authentication flow uses OAuth 2.0.") backend.write("mongodb_vfs/docs/overview.md", "This directory contains onboarding docs.") # Search for files that mention "authentication flow" # Newly written files can take a few seconds to become searchable result = backend.grep("authentication flow", path="mongodb_vfs/docs/") print("grep matches:") for match in result.matches or []: print(match["path"], match["line"], match["text"]) # Find files that match a glob pattern result = backend.glob("*.txt", path="mongodb_vfs/docs/") print("glob matches:", result.matches) # List the contents of a directory result = backend.ls("mongodb_vfs/docs/") print("ls entries:", result.entries) print("init_errors:", backend.init_errors)
デフォルトでは 、バックエンドはすべての操作をバケット内の mongodb_vfs/ プレフィックスに制限します。これを変更するには別の s3_prefix 値を MongoFilesystemBackend に渡します。または、バケット全体にアクセスするには s3_prefix="" を渡します。
注意
このバックエンドをDepAgentsエージェントに接続する方法については、 「DeepAgents クイックスタート」 を参照してください。
MongoDBエージェント ツールキー
MongoDB Agent Tools は、LorgGraph React Agent に渡すと、 MongoDBリソースを操作できるようにするツールのコレクションです。
利用可能なツール
名前 | 説明 |
|---|---|
| MongoDBデータベース をクエリするためのツール。 |
| MongoDBデータベースに関するメタデータを取得するためのツール。 |
| MongoDB database のコレクション名を取得するためのツール。 |
| LM を呼び出してデータベースクエリが正しいかどうかを確認するツール。 |
使用法
from langchain_openai import ChatOpenAI from langgraph.prebuilt import create_react_agent from langchain_mongodb.agent_toolkit import ( MONGODB_AGENT_SYSTEM_PROMPT, MongoDBDatabase, MongoDBDatabaseToolkit, ) db_wrapper = MongoDBDatabase.from_connection_string( "<connection-string>", database="<database-name>" ) llm = ChatOpenAI(model="gpt-4o-mini", timeout=60) toolkit = MongoDBDatabaseToolkit(db=db_wrapper, llm=llm) system_message = MONGODB_AGENT_SYSTEM_PROMPT.format(top_k=5) test_query = "Which country's customers spent the most?" agent = create_react_agent(llm, toolkit.get_tools(), prompt=system_message) agent.step_timeout = 60 events = agent.stream( {"messages": [("user", test_query)]}, stream_mode="values", ) messages = [] for event in events: messages.extend(event["messages"]) print(messages[-1].content)
注意
ドキュメントローダー
ドキュメントローダーは LangChain アプリケーションにデータをロードするのに役立つツールです。
MongoDBLoader は、MongoDB データベースからドキュメントのリストを返すドキュメントローダーです。
使用法
from langchain_mongodb.loaders import MongoDBLoader loader = MongoDBLoader.from_connection_string( connection_string="<connection-string>", # MongoDB cluster URI db_name="langchain_db", # Database that contains the collection collection_name="documents", # Collection to load documents from filter_criteria={"category": "ai"}, # Optional document to specify a filter field_names=["title", "summary"], # Optional list of fields to include metadata_names=["category"], # Optional metadata fields to extract ) docs = loader.load()
注意
チャット履歴
MongoDBChatMessageHistory は、 MongoDBデータベースにチャット メッセージ履歴を保存および管理できるコンポーネントです。一意のセッション識別子に関連付けられているユーザー メッセージとAI が生成したメッセージの両方を保存できます。このコンポーネントは、チャットボットなど、時間の経過とともにインタラクションを追跡するアプリケーションに使用します。
使用法
from langchain_mongodb.chat_message_histories import MongoDBChatMessageHistory chat_message_history = MongoDBChatMessageHistory( session_id="<session-id>", # Unique session identifier connection_string="<connection-string>", # MongoDB cluster URI database_name="langchain_db", # Database to store the chat history collection_name="chat_history", # Collection to store the chat history ) chat_message_history.add_user_message("Hello") chat_message_history.add_ai_message("Hi")
print(chat_message_history.messages)
[HumanMessage(content='Hello', additional_kwargs={}, response_metadata={}), AIMessage(content='Hi', additional_kwargs={}, response_metadata={}, tool_calls=[], invalid_tool_calls=[])]
ストレージ
MongoDB でデータを管理および保存するために、次のカスタム データ ストアを使用できます。
ドキュメント ストア
MongoDBDocStore は、MongoDB を使用してドキュメントを保存および管理するカスタム キーバリュー ストアです。CRUD 操作は、他の MongoDB コレクションと同様に実行できます。
使用法
from langchain_mongodb.docstores import MongoDBDocStore # Replace with your MongoDB connection string and namespace connection_string = "<connection-string>" namespace = "<database-name>.<collection-name>" # Initialize the MongoDBDocStore docstore = MongoDBDocStore.from_connection_string(connection_string, namespace)
注意
バイナリストレージ
MongoDBByteStore は、MongoDB を使用してバイナリデータ、具体的にはバイトで表されるデータを保存および管理するカスタム データストアです。キーが文字列で値がバイト シーケンスであるキーと値のペアを使用して、CRUD 操作を実行できます。
使用法
from langchain_community.storage.mongodb import MongoDBByteStore # Instantiate the MongoDBByteStore mongodb_store = MongoDBByteStore( connection_string="<connection-string>", # MongoDB cluster URI db_name="langchain_db", # Name of the database collection_name="byte_store", # Name of the collection ) # Set values for keys mongodb_store.mset([("key1", b"hello"), ("key2", b"world")]) # Get values for keys values = mongodb_store.mget(["key1", "key2"]) print(values) # Iterate over keys for key in mongodb_store.yield_keys(): print(key) # Delete keys mongodb_store.mdelete(["key1", "key2"])
[b'hello', b'world'] key1 key2
注意
追加リソース
MongoDBと LgGraph を統合する方法については、「 MongoDBと LgGraph の統合 」を参照してください。
インタラクティブPythonノートについては、 Docs Notes リポジトリ およびジェネレーティブAIが使用するリポジトリ を参照してください。