AI エージェント向け: ドキュメントインデックスは https://www.mongodb.com/ja-jp/docs/llms.txt で利用できます。すべてのページの markdown バージョンは、いずれかの URL パスに .md を追加することで利用できます。
Docs Menu

言語アナライザ

言語固有のアナライザ を使用して、特定の言語にカスタマイズされたインデックスを作成します。 各言語アナライザには、その言語の使用パターンに基づくストップワードと単語の除算が組み込まれています。

MongoDB Search は、次の言語アナライザを提供します。

lucene.arabic

lucene.armenian

lucene.basque

lucene.bengali

lucene.brazilian

lucene.bulgarian

lucene.catalan

lucene.chinese

lucene.cjk 1

lucene.czech

lucene.danish

lucene.dutch

lucene.english

lucene.finnish

lucene.french

lucene.galician

lucene.german

lucene.greek

lucene.hindi

lucene.hungarian

lucene.indonesian

lucene.irish

lucene.italian

lucene.japanese

lucene.korean

lucene.kuromoji 2

lucene.latvian

lucene.lithuanian

lucene.morfologik 3

lucene.nori 4

lucene.norwegian

lucene.persian

lucene.polish

lucene.portuguese

lucene.romanian

lucene.russian

lucene.smartcn 5

lucene.sorani

lucene.spanish

lucene.swedish

lucene.thai

lucene.turkish

lucene.ukrainian

1 cjk は、一般的な中国語、日本語、大文字と小文字のアナライザです

2 kuromoji は、日本語のアナライザです

3 morfologik は、ポーランド語アナライザです

4 nori は、 韓国語アナライザ

5 smartcn は中国語アナライザです

次のドキュメントを含むcarsという名前のコレクションについて考えてみます。

{
"_id": 1,
"subject": {
"en": "It is better to equip our cars to understand the causes of the accident.",
"fr": "Mieux équiper nos voitures pour comprendre les causes d'un accident.",
"he": "עדיף לצייד את המכוניות שלנו כדי להבין את הגורמים לתאונה."
}
}
{
"_id": 2,
"subject": {
"en": "The best time to do this is immediately after you've filled up with fuel",
"fr": "Le meilleur moment pour le faire c'est immédiatement après que vous aurez fait le plein de carburant.",
"he": "הזמן הטוב ביותר לעשות זאת הוא מיד לאחר שמילאת דלק."
}
}

次のインデックス定義の例では、 frenchアナライザを使用してsubject.frフィールドのインデックスを指定します。

{
"mappings": {
"fields": {
"subject": {
"fields": {
"fr": {
"analyzer": "lucene.french",
"type": "string"
}
},
"type": "document"
}
}
}
}

The following MongoDB Search query searches for the string pour in the subject.fr field. To run this query, connect to your cluster using mongosh and switch to the database that contains the cars collection.

db.cars.aggregate([
{
$search: {
"text": {
"query": "pour",
"path": "subject.fr"
}
}
},
{
$project: {
"_id": 0,
"subject.fr": 1
}
}
])

上記のクエリでは、 frenchアナライザを使用しても結果が返されません。 pourは組み込みのストップワードであるためです。 standardアナライザを使用すると、同じクエリで両方のドキュメントが返されます。

The following MongoDB Search query searches for the string carburant in the subject.fr field. To run this query, connect to your cluster using mongosh and switch to the database that contains the cars collection.

db.cars.aggregate([
{
$search: {
"text": {
"query": "carburant",
"path": "subject.fr"
}
}
},
{
$project: {
"_id": 0,
"subject.fr": 1
}
}
])

MongoDB Search では、結果に _id: 1 を含むドキュメントが返されます。このドキュメントは、ドキュメント用に lucene.frenchアナライザが作成したトークンとクエリが一致したためです。lucene.frenchアナライザは、_id: 1 を使用してドキュメントの subject.frフィールドに次のトークンを作成します。

meileu

moment

fair

est

imediat

aprè

fait

plein

carburant

また、 icFlexing と ストップワード トークン フィルターを使用して カスタムアナライザ を作成し、サポートされていない言語のインデックスを作成することもできます。

次のインデックス定義の例では、 myHebrewAnalyzerというカスタムアナライザを使用してヘブライ テキストのトークンを分析および作成し、 subject.heフィールドのインデックスを指定します。

{
"analyzer": "lucene.standard",
"mappings": {
"dynamic": false,
"fields": {
"subject": {
"fields": {
"he": {
"analyzer": "myHebrewAnalyzer",
"type": "string"
}
},
"type": "document"
}
}
},
"analyzers": [
{
"charFilters": [],
"name": "myHebrewAnalyzer",
"tokenFilters": [
{
"type": "icuFolding"
},
{
"tokens": [
"אן",
"שלנו",
"זה",
"אל"
],
"type": "stopword"
}
],
"tokenizer": {
"type": "standard"
}
}
]
}

The following MongoDB Search query searches for the string המכוניות in the subject.he field. To run this query, connect to your cluster using mongosh and switch to the database that contains the cars collection.

db.cars.aggregate([
{
$search: {
"text": {
"query": "המכוניות",
"path": "subject.he"
}
}
},
{
$project: {
"_id": 0,
"subject.he": 1
}
}
])

MongoDB Search では、結果に _id: 1 を含むドキュメントが返されます。このクエリは、myHebrewAnalyzerアナライザがドキュメント用に作成したトークンとクエリが一致したためです。myHebrewAnalyzerアナライザは、_id: 1 を使用してドキュメントの subject.heフィールドに次のトークンを作成します。

עדיף

לצייד

את

המכוניות

כדי

להבין

את

הגורמים

לתאונה

複数の言語アナライザを使用して多言語検索を実行するインデックスを作成することも可能です。

次のインデックス定義の例では、 sample_mflix.moviesコレクションに 動的マッピング を含むインデックスを指定しています。この定義は、lucene.italian言語アナライザを適用して fullplotフィールドをインデックスし、マルチオプションを使用して lucene.english を代替言語アナライザとして指定します。MongoDB Search は、 moviesコレクション内の動的にインデックスを作成する他のすべてのフィールドに対してデフォルトのlucene.english言語アナライザを使用します。

{
"analyzer": "lucene.standard",
"mappings": {
"dynamic": true,
"fields": {
"fullplot": {
"type": "string",
"analyzer": "lucene.italian",
"multi": {
"fullplot_english": {
"type": "string",
"analyzer": "lucene.english",
}
}
}
}
}
}

The following MongoDB Search query uses the compound operator to query the collection in multiple languages. To run this query, connect to your cluster using mongosh and switch to the sample_mflix database.

複合演算子には、次の句が含まれます。

  • must 句は、text 演算子を使用して Bella という用語を含む英語とイタリア語の映画のプロットを検索します

  • mustNot 句は、1984 年から 2016 年の間に公開された映画を range 演算子を使用して除外します

  • should 句は、text 演算子を使って Comedy ジャンルの好みを指定します

db.movies.aggregate([
{
$search: {
"index": "multilingual-tutorial",
"compound": {
"must": [{
"text": {
"query": "Bella",
"path": { "value": "fullplot", "multi": "fullplot_english" }
}
}],
"mustNot": [{
"range": {
"path": "released",
"gt": ISODate("1984-01-01T00:00:00.000Z"),
"lt": ISODate("2016-01-01T00:00:00.000Z")
}
}],
"should": [{
"text": {
"query": "Comedy",
"path": "genres"
}
}]
}
}
},
{
$project: {
"_id": 0,
"title": 1,
"plot": 1,
"genres": 1,
"runtime": 1,
"fullplot": 1,
"released": 1,
"score": { "$meta": "searchScore" }
}
}
])