You can explore and deploy Voyage AI by MongoDB models from Model Garden on Gemini Enterprise Agent Platform, formally GCP Vertex AI.
모델 가든은 MongoDB 모델을 통한 보야지 AI 의 라이선스를 관리하고, 온디맨드 hardware 또는 기존 Compute Engine 예약을 사용하여 배포서버 옵션을 제공합니다.
Voyage AI by MongoDB models are self-deployed partner models, meaning you pay for both the model usage and the Gemini Enterprise Agent Platform infrastructure consumed. Gemini Enterprise Agent Platform handles deployment and provides endpoint management features.
사용 가능한 모델
To see which models you can deploy, search for "Voyage" in the Model Garden on Gemini Enterprise Agent Platform.
Voyage AI 모델에 대해 자세히 학습하려면 모델 개요를 참조하세요.
가격
Pricing for Voyage AI by MongoDB models in Model Garden on Gemini Enterprise Agent Platform includes:
모델 사용 요금: 시간당 요금으로 청구되는 Voyage AI 모델 컨테이너 사용 비용 입니다. 사용 요금은 배포서버 위해 선택한 특정 모델 및 hardware 구성에 따라 달라집니다. 자세한 가격 정보는 Google Cloud Marketplace 모델 목록 페이지의 가격 섹션을 참조하세요.
해당 리전의 Google Cloud 기본 인스턴스: 특정 리전에 특정한 기본 Google Cloud GPU 인스턴스 (예: N4, A100 또는 H100)의 비용은 월 단위로 청구되며 다음 기준에 따라 가격이 책정됩니다. vCPU. 자세한 학습 은 Google Cloud Compute Engine 가격을 참조하세요.
All billing charges appear as the use of Gemini Enterprise Agent Platform on your Google Cloud bill.
특정 Voyage AI 모델의 가격을 보려면 다음 단계를 따르세요.
Quotas
When you deploy Voyage AI models, you consume Gemini Enterprise Agent Platform resources that are subject to quotas. You can view and manage your quotas in the Quotas section of the Google Cloud Console's IAM page. For more information, see View the quotas for your project. In the same page, you can right-click any current quota, click Edit quota, and submit a request to increase your quota if needed.
전제 조건
To get started using the Voyage AI by MongoDB models through Gemini Enterprise Agent Platform, you must:
Google Cloud 프로젝트 및 개발 환경을 설정합니다. 지침은 프로젝트 및 개발 환경 설정을 참조하세요.
Enable the Gemini Enterprise Agent Platform API. For instructions, see Setup.
하드웨어 구성
Each model in the Model Garden lists its recommended hardware configuration. Consult Model Garden on Gemini Enterprise Agent Platform for each Voyage model's recommended hardware specifications.
For example, for the voyage-4 model, use the following recommended instances that Model Garden on Gemini Enterprise Agent Platform suggests for deployment. These recommendations may change and we recommend that you consult the official Model Garden on Gemini Enterprise Agent Platform page for a particular Voyage AI model to see its recommended hardware.
A2 인스턴스(예:
a2-highgpu-1g또는a2-ultragpu-1g)와 A100 GPU가 기본값 선택됩니다.더 높은 성능 요구 사항에는3 H GPU를 사용하는
a3-highgpu-1g와 같은 A 인스턴스를 사용하는 것이 좋습니다.100
지원되는 리전
모델 가든에는 각 Voyage AI 모델에 대해 지원되는 리전이 나열되어 있습니다. 다른 리전에서 모델에 대해 지원이 필요한 경우 MongoDB 지원문의.
권장사항 및 제한 사항
엔드포인트 유형: 모든 Voyage AI 모델에는 전용 공개 엔드포인트 유형이 필요합니다. 자세한 내용은 엔드포인트 유형 선택을 참조하세요.
input_type 이해: 쿼리 대 문서:
input_type매개 변수는 검색 작업을 위한 임베딩을 최적화합니다. 검색 쿼리에는"query"를 사용하고 검색 중인 콘텐츠에는"document"를 사용합니다. 이 최적화는 검색 정확도를 향상시킵니다.input_type매개변수에 대해 자세히 알아보려면 임베딩 및 재지정 API 개요를 참조하세요.다른 출력 차원 사용: Voyage 4 모델은 256, 512, 1024 (기본값), 2048 등 여러 출력 차원을 지원. 차원이 작을수록 저장 및 계산 비용이 줄어들고, 차원이 클수록 정확도가 향상될 수 있습니다. 정확도 요구 사항과 리소스 제약 조건의 균형을 가장 잘 맞추는 차원을 선택하세요.
Voyage AI 모델 찾기
모델 가든에서 MongoDB 모델의 Voyage AI 찾으려면 다음을 수행합니다.
Voyage 모델을 검색합니다.
Search Models 필드 에 "Voyage"를 입력하여 MongoDB 모델의Voyage AI 목록을 표시합니다.
참고
The Google Cloud Marketplace has two search boxes: one for the entire Marketplace and one within Model Garden on Gemini Enterprise Agent Platform. To locate Voyage AI by MongoDB models, use the search box on Model Garden on Gemini Enterprise Agent Platform.
또는 Model Garden > Model Collections > Partner Models를 통해 Voyage AI 모델로 이동한 다음 여기에 나열된 Voyage AI 모델 중 하나를 선택할 수 있습니다.
Task-specific solutions 까지 아래로 스크롤하여 있는 그대로 사용하거나 필요에 맞게 사용자 지정할 수 있는 Voyage AI 모델을 찾을 수도 있습니다.
Deploy a Voyage AI Model in Gemini Enterprise Agent Platform
MongoDB 의 Voyage AI 모델을 사용하여 예측하려면 온라인 추론을 위해 이를 비공개 전용 엔드포인트에 배포 해야 합니다. 배포서버는 지연 시간이 짧고 처리량이 높은 온라인 예측을 위해 물리적 리소스를 모델과 연결합니다. 여러 모델을 하나의 엔드포인트에 배포 하거나 동일한 모델을 여러 엔드포인트에 배포할 수 있습니다.
모델을 배포 때 다음 옵션을 고려하세요.
엔드포인트 위치
모델 컨테이너
모델 실행 에 필요한 컴퓨팅 리소스
모델을 배포 후에는 이러한 설정을 변경할 수 없습니다. 배포서버 구성을 수정해야 하는 경우 모델의 배포를 취소하고 새 설정으로 다시 배포해야 합니다.
Voyage AI models require a dedicated public endpoint. For more information, see Create a public endpoint in the Gemini Enterprise Agent Platform documentation.
To deploy a model in Gemini Enterprise Agent Platform using the console:
모델을 찾습니다.
모델 가든 콘솔 고 (Go) Search Models 필드 에서 "Voyage"를 검색 MongoDB 모델의 Voyage AI 목록을 표시합니다.
모델을 활성화하고 계약에 동의합니다.
Enable를 클릭합니다. MongoDB Marketplace 최종 사용자 계약 이 열립니다. 계약을 검토하고 동의하여 모델을 활성화 하고 필요한 상업용 라이선스를 받습니다.
배포서버 옵션을 검토합니다.
계약에 동의하면 모델 페이지에 다음 옵션이 표시됩니다.
Deploy a model: 모델을 모델 레지스트리에 저장하고 Google Cloud의 엔드포인트에 배포합니다. 콘솔을 사용하여 배포 하려면 다음 단계를 계속 진행합니다.
Create an Open Notebook for Voyage Embedding Models Family: 협업 환경에서 모델을 미세 조정 및 사용자 지정하고 최적의 비용 과 성능을 위해 모델을 혼합할 수 있습니다. Voyage AI 용 Vertex AI 노트북 샘플을 참조하세요.
View Code: 모델 배포 및 사용을 위한 코드 샘플을 표시합니다. 코드를 사용하여 프로그래밍 방식으로 배포 하려면 코드를 사용하여 배포를 참조하세요.
배포서버 양식을 작성합니다.
A form opens that allows you to review and edit the deployment options. Gemini Enterprise Agent Platform provides default settings that are optimized for the model, but you can customize them as needed. For example, you can select the machine type, GPU type, and number of replicas. The following example shows default settings for the voyage-4 model, but these may change, so review the settings carefully before deploying.
필드 | 설명 |
|---|---|
Resource ID | 드롭다운 메뉴에서 선택합니다(미리 선택됨). |
Model Name | 드롭다운 메뉴에서 선택합니다(미리 선택됨). |
Region | 원하는 리전(예: |
Endpoint name | 엔드포인트의 이름(예: |
Serving spec | 머신 유형(예: |
Accelerator type |
|
Accelerator count |
|
Replica count | 복제본의 최소 및 최대 수(예: |
Reservation type | 예약 유형(예: |
VM provisioning model | 프로비저닝 모델(예: |
Endpoint access | Public (Dedicated endpoint)0}을 선택합니다. |
코드를 사용하여 배포
모델 세부 정보 페이지에서 View Code 를 선택한 경우 Vertex AI SDK를 사용하여 프로그래밍 방식으로 모델을 배포 할 수 있습니다. 이 접근 방식은 코드를 통해 배포서버 구성을 완전히 제어할 수 있습니다.
Google Cloud Vertex AI SDK에 대한 자세한 내용은 Python 용 Vertex AI SDK 설명서를 참조하세요.
참고
이 섹션의 코드 예시는 voyage-4 모델용이며 변경될 수 있습니다. 최신 코드 예제는 모델 가든의 모델 페이지에 있는 View Code 탭 참조하세요. 다른Voyage AI 모델의 경우 코드가 비슷하지만 모델별 세부 정보는 Model Garden에서 해당 모델의 페이지를 확인하세요.
코드를 사용하여 모델을 배포 하려면 다음을 수행합니다.
엔드포인트에 배포합니다.
새 모델을 배포 할지, 아니면 기존 엔드포인트를 사용할지 선택합니다.
# Choose whether to deploy a new model or use an existing endpoint: deployment_option = "deploy_new" # ["deploy_new", "use_existing"] # If using existing endpoint, provide the endpoint ID: ENDPOINT_ID = "" # {type:"string"} if deployment_option == "deploy_new": print("Deploying new model...") endpoint = model.deploy( machine_type="a3-highgpu-1g", accelerator_type="NVIDIA_H100_80GB", accelerator_count=1, accept_eula=True, use_dedicated_endpoint=True, ) print(f"Endpoint deployed: {endpoint.display_name}") print(f"Endpoint resource name: {endpoint.resource_name}") else: if not ENDPOINT_ID: raise ValueError("Please provide an ENDPOINT_ID when using existing endpoint") from google.cloud import aiplatform print(f"Connecting to existing endpoint: {ENDPOINT_ID}") endpoint = aiplatform.Endpoint( endpoint_name=f"projects/{PROJECT_ID}/locations/{LOCATION}/endpoints/{ENDPOINT_ID}" ) print(f"Using endpoint: {endpoint.display_name}") print(f"Endpoint resource name: {endpoint.resource_name}")
중요
Voyage AI 모델에는 전용 공개 엔드포인트가 필요하므로 use_dedicated_endpoint 를 True 로 설정합니다.
Gemini Enterprise Agent Platform deploys the model to a managed endpoint that you can access to make online inferences or batch inferences through the Google Cloud console or the Gemini Enterprise Agent Platform API.
For more information, see Deploy a model to an endpoint in the Gemini Enterprise Agent Platform documentation.
예측합니다.
After deployment, you can make predictions using the Gemini Enterprise Agent Platform endpoint.
모든 엔드포인트 매개변수 및 예측 옵션에 대해서는 임베딩 및 API 재순위 지정 개요를 참조하세요.
import json # Multiple texts to embed texts = [ "Machine learning enables computers to learn from data.", "Natural language processing helps computers understand human language.", "Computer vision allows machines to interpret visual information.", "Deep learning uses neural networks with multiple layers." ] # Prepare the batch request and make invoke call body = { "input": texts, "output_dimension": 1024, "input_type": "document" } response = endpoint.invoke( request_path="/embeddings", body=json.dumps(body).encode("utf-8"), headers={"Content-Type": "application/json"} ) # Extract embeddings result = response.json() embeddings = [item["embedding"] for item in result["data"]] print(f"Number of texts embedded: {len(embeddings)}") print(f"Embedding dimension: {len(embeddings[0])}") print(f"\nFirst embedding (first 5 values): {embeddings[0][:5]}") print(f"Second embedding (first 5 values): {embeddings[1][:5]}")
모델 배포 취소 및 엔드포인트 삭제
배포된 모델과 해당 엔드포인트를 제거 하려면 다음을 수행합니다.
엔드포인트에서 모델 배포를 취소합니다.
선택적으로 엔드포인트 자체를 삭제 .
For detailed instructions, see Undeploy a model and delete the endpoint in the Gemini Enterprise Agent Platform documentation.
중요
엔드포인트에서 모든 모델의 배포가 취소된 후에만 엔드포인트를 삭제 수 있습니다. 모델 배포를 취소하고 엔드포인트를 삭제하면 해당 엔드포인트에 대한 모든 추론 서비스 및 청구가 중지됩니다.