Vector Quantization in Atlas Vector Search Now Generally Available
April 30, 2025
What it is:
Vector quantization in Atlas Vector Search enables more efficient vector storage and ingestion with automatic, built-in processing. Developers can choose scalar or binary quantization with rescoring, reducing memory usage by up to 96% while maintaining high retrieval accuracy.
Who is it for:
These vector quantization capabilities are built for teams scaling AI, semantic search, and retrieval-augmented generation (RAG) workloads. This capability combines MongoDB’s flexible document model with optimized vector compression, delivering lower cost, faster performance, and easier scalability for advanced use cases.
Why it matters:
By compressing vector data with minimal accuracy loss, vector quantization makes search and AI workloads more cost-efficient and performant. Running natively in Atlas eliminates the need for external preprocessing pipelines, helping developers simplify deployments and scale confidently across environments.
How to get started:
Click the documentation link below for setup guidance, configuration options, and best practices for vector quantization usage in Atlas Vector Search.
Related Content
Vector Quantization: Scale Search & Generative AI Applications