Category report

Vector databases and nearest-neighbor search libraries

Research date: 2026-10-09.

Scope: 28 GitHub repositories spanning distributed vector databases, persistent vector indexes, SQL extensions, CPU/GPU approximate nearest-neighbor search, locality-sensitive hashing, non-metric search, and exact spatial nearest-neighbor libraries. The selection emphasizes implementations worth studying, rather than hosted products, client SDKs, embedding models, or benchmark rankings. Each repository heading links to its verified canonical GitHub location; linked documentation and implementation files provide starting points for study.

Criteria legend: C1 — difficult correctness involving invariants, concurrency, numerical semantics, adversarial inputs, or failures. C2 — substantial reusable abstractions supporting multiple use cases. C3 — concrete performance constraints addressed through an understandable architecture. C4 — sustained evolution with evidence of compatibility, testing, or complexity management. Criterion assignments are grounded engineering judgments about the cited material, not certifications of correctness or uniform code quality.

Distributed vector databases and services

1. qdrant/qdrant

Rust — vector database with payload filtering and segmented storage. Study how a mutable database surrounds similarity indexes with identity mapping, persistence, and storage policy.

  • C1: Operations receive WAL sequence numbers, and segments and individual points retain versions so replay can reject stale updates. Search also deduplicates points that occur in multiple segments. These are concrete recovery and identity invariants. Storage and versioning.
  • C2, C3: A segment combines vector/payload stores, corresponding indexes, and internal/external ID mapping. Appendable versus non-appendable segments and configurable memory placement separate mutation requirements from query cost. Indexed payload values can avoid repeated disk reads during filtering. The same storage guide explains these boundaries and tradeoffs.

2. milvus-io/milvus

Go and C++ — distributed vector database. A useful study of the machinery needed to serve both newly ingested vectors and large, indexed historical datasets.

  • C1: The documented insertion path records writes in the WAL before processing; streaming nodes coordinate shard recovery and the transition from growing to sealed data. Search combines results from growing data and historical segments, requiring consistent routing and reduction. Architecture and data flows.
  • C2, C3: Stateless proxies, a coordinator, streaming/query/data workers, metadata storage, and object storage have distinct responsibilities. Compaction and index construction run separately from historical queries, while results are reduced at several levels. This makes compute/storage separation and independent scaling tangible. The architecture page is the principal entry point; its current description is version-sensitive and should not be assumed to describe older Milvus deployments.

3. weaviate/weaviate

Go — object and vector database with inverted indexes. Study why different indexes in one shard use different persistence structures.

  • C1: Objects and vectors can change while reads continue; the storage design includes WAL-based persistence and HNSW snapshots. Maintaining correspondence among object, inverted, and vector stores creates useful consistency and recovery problems to investigate. Storage design.
  • C2, C3: Each shard packages an object store, inverted index, and vector index. Object/inverted data use LSM segments, Bloom filters, and background merging; the vector graph is kept outside that segmentation because merging HNSW graphs is costly. The explicit explanation of this asymmetry makes Weaviate particularly instructive for storage-engine design.

4. chroma-core/chroma

Rust and Python — vector, full-text, and metadata search infrastructure. Focus on the distributed execution and storage implementation, rather than just the Python collection API.

  • C1: Writes are acknowledged after log persistence; query executors consult the log alongside materialized indexes, while compactors publish new index versions through the system database. This exposes the consistency problem between fresh writes and asynchronously built indexes. Distributed architecture.
  • C2, C3: Gateway, log, query executor, compactor, and catalog are separate components. Collection-based rendezvous routing and local SSD caches address the latency of object storage. Study these boundaries in the Rust source workspace. Deployment-specific behavior matters: the cited design describes distributed Chroma, not every local configuration.

5. vdaas/vald

Go, integrating native ANN engines — Kubernetes-oriented distributed vector search. Vald is included for its distributed serving implementation; NGT, listed separately below, supplies a different layer of the stack.

  • C2: Agents, gateways, discovery, and index management form a reusable service architecture around vector indexes. Components can be deployed selectively, with Kubernetes integration handling discovery and lifecycle concerns. Architecture.
  • C3: Vectors and indexes are distributed among agents; queries execute in parallel and the load-balancing gateway merges results. This is a concrete fan-out/fan-in design to study alongside index placement and backup restoration. Its architectural value lies in coordinating search services rather than introducing another nearest-neighbor algorithm.

SQL integration and persistent index storage

6. pgvector/pgvector

C — PostgreSQL vector types and HNSW/IVFFlat access methods. An especially useful bridge between ANN algorithms and a transactional database executor.

  • C1: HNSW scanning requires MVCC snapshots, coordinates with vacuum through page locks, checks vector dimensions, and accounts for nodes being deleted or replaced between iterative scans. HNSW scan implementation.
  • C3: Iterative search resumes from discarded candidates and respects tuple and memory budgets derived from PostgreSQL settings. Strict ordering has an explicit distance check; these details reveal how an approximate index cooperates with SQL execution.
  • C4: The dated changelog documents multiple years of PostgreSQL compatibility changes, parallel-build improvements, vacuum/concurrency fixes, and memory-safety repairs. This is evidence of sustained complexity management, not an assertion that these failure modes never existed.

7. supervc-stack/VectorChord

Rust — PostgreSQL extension for quantized, partitioned vector search. The former tensorchord/VectorChord URL redirects here. Study an alternative to keeping an entire proximity graph in memory.

  • C2: Its vchordrq access method works through PostgreSQL operator classes for several vector representations and distance semantics, reusing pgvector-compatible SQL types and query syntax. Indexing guide.
  • C3: Building separates vector-space partitioning from row insertion, with threads for the former and PostgreSQL processes for the latter. Query probes select lists; residual quantization and spherical centroids are explicit tuning choices. RaBitQ compression and reranking are described in the repository overview. Together these provide a concrete study of memory, build-time, and recall tradeoffs without relying on its advertised cost or throughput comparisons.

8. asg017/sqlite-vec

C — portable SQLite vector-search extension. A compact implementation of embedded KNN search with meaningful query-planner and storage decisions.

  • C1: Metadata predicates participate in KNN calculation rather than merely filtering a completed result set. The documentation specifies strict metadata types, supported operators, and limitations of auxiliary columns; unsupported predicates cannot be assumed to preserve the intended result semantics. The vec0 virtual table.
  • C2, C3: Metadata columns, partition keys, and auxiliary columns deliberately have different physical and query roles. Partition keys colocate vectors and restrict scans, while auxiliary data live in a separate internal table. The guide explains the risk of oversharding. This is valuable for studying exact scan-oriented search inside SQLite, rather than assuming every vector extension is an ANN graph.

9. lance-format/lance

Rust with Python interfaces — columnar/table storage with a substantive vector-index subsystem. The former lancedb/lance URL redirects here. This entry covers Lance's index/storage implementation, not the separately packaged LanceDB database.

  • C2: Vector indexes explicitly compose clustering, a sub-index, and quantization: IVF can combine with flat or HNSW search and several quantizers. Vector index specification.
  • C1, C3: Index structure and quantized vectors occupy separate Lance files with partition offsets, codebooks, and versioned metadata. Partition ordering and HNSW-level ordering are stated invariants. Readers must validate carried-column field identities before serving covering values. These details make the format a useful study of efficient access coupled to safe interpretation of persistent data. The repository also distinguishes stable file-format compatibility from SDK/API compatibility.

Reusable graph, quantization, and disk-search libraries

10. facebookresearch/faiss

C++, GPU kernels, and Python — broad vector-search and quantization library. Study how a common index model accommodates exact search, graphs, inverted files, compression, and refinement.

  • C2: Index implementations compose coarse quantizers, encoded vector storage, graph traversal, and refinement stages. HNSW can use flat, scalar-quantized, product-quantized, or two-level storage. Index architecture guide.
  • C3: IVF restricts candidate lists, PQ reduces stored representation, and residual quantization uses beam search to balance encoding accuracy against work. The guide explains distinct memory/search/training controls and constraints, including the absence of deletion support in its HNSW implementation. Faiss is particularly useful for comparing architectures behind one reusable interface; an index's supported operations must be checked individually.

11. nmslib/hnswlib

C++ with Python bindings — compact HNSW implementation. A focused entry point for understanding graph construction and search before exploring a full database.

  • C1: The API states a precise concurrency boundary: concurrent insertions are supported, and concurrent queries are supported, but inserting while querying is not. Resizing and serialization add further restrictions; deleted-label handling also has explicit semantics. These are useful invariants to trace in a small codebase.
  • C3: M controls graph connectivity and memory, ef_construction controls construction exploration, and ef controls query exploration. ef must accommodate the requested result count, and insufficient reachable results can cause an exception. Algorithm parameters. The repository's API documentation and this parameter guide are the two complementary entry points.

12. nmslib/nmslib

C++ with Python bindings — generic metric and non-metric similarity search. Retained separately from hnswlib because its space/method framework and evaluation machinery are substantially broader.

  • C1: Search spaces can violate the triangle inequality and can be asymmetric; left-versus-right query orientation therefore matters. This makes it useful for studying assumptions often hidden by Euclidean-only APIs. Problem formulation and limitations.
  • C2: Methods and spaces have separate factories and distance-value templates. RangeQuery and KNNQuery share radius/result-update operations, allowing algorithms to reuse pruning logic. The extension guide explains this design and how its benchmark executor generates gold-standard results. Persistence, updates, and range search are method-dependent rather than uniform capabilities.

13. spotify/annoy

C++ with Python bindings — random-projection-tree ANN indexes. Study an architecture built around immutable, distributable search artifacts.

  • C2: Building and serving are separate phases. A built index can be serialized and mapped by multiple processes; several distance implementations share the same index API. The repository explains the restriction that additional items cannot be appended after construction.
  • C3: Memory-mapped static indexes allow shared pages across workers, and optional prefaulting trades loading work for later page-fault behavior. Tree count and the query's node-exploration budget are separate accuracy/resource controls. Core implementation exposes the node layout, mapping helpers, and distance routines. It is a useful static-index design reference, not evidence of transactional durability or online mutation support.

14. microsoft/DiskANN

Rust in the current mainline; historical C++ implementation retained on cpp_main — composable vector indexing. The mainline now describes DiskANN3, so older C++ tutorials do not describe its present architecture.

  • C2: A DataProvider trait separates index algorithms from storage and retrieval of vectors and adjacency lists. Example providers target memory, disk, a key-value engine, and a tree-based store. The repository explains this contract and its filtering/update interfaces.
  • C3: Memory tiers, quantizers, and provider-specific storage paths make the cost of graph traversal an explicit integration concern. The core crate documentation also describes serialized algorithm-test baselines and diagnostic comparisons, useful for understanding how behavior changes are reviewed.

Status distinction: the repository explicitly says the legacy C++ branch is no longer actively developed or maintained. It is not counted as a second project.

15. microsoft/SPTAG

C++ with language bindings — space-partition trees plus neighborhood graphs. A different graph-search design from an HNSW-only library.

  • C2: Its builder and searcher support both kd-tree/RNG and balanced-k-means-tree/RNG variants. Trees find starting seeds, after which tree and graph searches alternate. The repository explains why balanced k-means partitions are useful when kd-tree distance bounds become weak in high dimensions.
  • C3: Sample counts, graph neighborhood size, construction refinement, tree counts, and query node checks control distinct costs. The parameter guide explicitly maps controls to memory, build time, index quality, and recall/latency, and describes measurement against a truth set. This is a strong entry point for studying staged index construction and tuning.

16. unum-cloud/USearch

C++ core with numerous native bindings — extensible HNSW-based similarity indexing. Particularly useful for studying custom distance functions and externally stored objects.

  • C2: The core index accepts user-defined metrics, heterogeneous lookups, external values, and search predicates rather than requiring every application to use one dense-vector representation. Core index header.
  • C1: The same header states capacity-reservation requirements and a subtle update limitation: simultaneous updates to the same entry do not provide transactional guarantees and may leave values and neighbor lists inconsistent. This is a concrete distinction between thread-safe operations and transactional semantics.
  • C3: Compact identifiers, reduced-precision representations, caller-supplied execution, and memory-mapped serving address graph memory and deployment overhead. These mechanisms, rather than the README's comparative speed claims, justify inclusion.

17. NGT-labs/NGT

C++ — neighborhood graphs, tree seeding, and quantized search. The former yahoojapan/NGT URL redirects here.

  • C2: The implementation supports graph-only search, graph-plus-tree search, multiple graph variants, vector representations, and distance spaces. The command/algorithm guide explains how a tree selects graph entry points, while graph-only mode can choose random starts.
  • C3: Query edge limits and range coefficients adjust traversal work; graph reconstruction treats outgoing and incoming edges separately and includes shortcut reduction, prefetch tuning, and accuracy-table generation. Quantized representations provide another memory/accuracy axis. This makes NGT useful for comparing graph topology optimization with simple increases in search breadth. The graph implementation is a verified source entry point.

18. antgroup/vsag

C++ with Python and JavaScript/TypeScript bindings — vector-index library. Focus on HGraph and the separation of graph traversal from vector representation.

  • C2: HGraph exposes alternative graph construction methods, pluggable base quantizers, and an optional more precise representation for reranking behind one index interface. HGraph design and configuration.
  • C3: Upper layers provide navigation; the bottom layer has a degree budget, and search has a bounded candidate frontier. Optional reverse-edge tracking spends additional edge storage to accelerate reverse lookup, while two-representation search trades memory for score accuracy. These are specific architectural choices to compare with other graph indexes.
  • C1: Parameter compatibility and restoration requirements are documented, and the developer guide exposes AddressSanitizer and ThreadSanitizer test paths. Their availability is evidence of attention to failure modes, not proof that the implementation is defect-free.

19. datastax/jvector

Java — embedded graph search with disk-resident data. The former jbellis/jvector URL redirects here. Study graph algorithms under Java's memory and vectorization constraints.

  • C1: The repository describes nonblocking graph construction, while the upgrade guide records explicit ownership changes for readers and suppliers. It also distinguishes mutable and immutable compressed-vector containers.
  • C2, C3: An HNSW-like hierarchy uses Vamana within layers; upper layers reside in memory and the bottom layer can reside on disk. Two-pass search combines compressed scoring with more accurate reranking. The upgrade guide explains fused PQ, optional hierarchy, early termination, and associated disk-format requirements. This is a substantive Java implementation, not a binding to another ANN library.

Partitioned scoring, accelerators, and scientific graph construction

20. google-research/google-research — ScaNN subsystem

C++ with Python/TensorFlow interfaces — partitioned similarity search and anisotropic quantization. Only the ScaNN subtree is selected; the monorepo is counted once.

  • C2: ScannBuilder composes optional partitioning, candidate scoring, and optional more accurate rescoring. Brute-force scoring and asymmetric hashing fit the same pipeline. Algorithms and configuration.
  • C3: The guide separates partition count, searched partitions, quantization block size, and reranking breadth. It explains that quantized brute-force scoring helps memory-bandwidth-bound workloads but can lose that advantage when data fit in cache or batching changes the bottleneck. This is unusually useful evidence for studying when compression helps, rather than assuming compressed arithmetic is always faster.

21. NVIDIA/cuvs

C++/CUDA with several language interfaces — GPU vector search and clustering. Current canonical GitHub location, formerly encountered under rapidsai/cuvs in search results.

  • C2: The CAGRA implementation separates graph building, searching, and conversion to HNSW for CPU serving. It can initialize a graph using IVF-PQ or NN-Descent and then prune redundant edges. CAGRA implementation guide.
  • C1, C3: The guide exposes GPU execution strategies, working-set sizes, and device-memory requirements. Index extension has an explicit ownership contract: the caller supplies concatenated, padded vector storage, while extension grows the graph and rebinds its data view. This offers concrete lessons in GPU resource lifetime, construction scratch space, and the separation of indexing hardware from serving hardware. No cross-library benchmark advantage is assumed here.

22. ejaasaari/lorann

C++17 with Python bindings — low-rank regression-based ANN. A research-derived alternative to graph traversal that broadens the selection beyond HNSW and IVF-PQ.

  • C2: Templated data and query quantizers share a base interface with a floating-point variant. The API exposes cluster training, exact search, approximate search, and serialization. C++ reference.
  • C1, C3: Cluster count, globally reduced dimension, regression rank, training-neighborhood size, and reranking count have separate roles. Rank and dimension restrictions are documented rather than left implicit; approximate candidate scoring is followed by configurable reranking. The repository identifies CPU instruction-set optimizations and labels GPU support experimental. Study how low-rank score approximation changes arithmetic and storage costs; its published speed claims are not treated as universal results.

23. lmcinnes/pynndescent

Python with compiled numerical execution — NN-Descent graph construction and querying. Useful both as a search library and as an input to manifold-learning or other neighbor-graph algorithms.

  • C2: It supports numerous distance families, custom distances, and a scikit-learn-compatible neighbor transformer. The graph itself is a useful reusable output, not merely a hidden serving structure.
  • C3: Random-projection trees provide initial neighbors; iterative neighbor-of-neighbor refinement improves the graph; diversification and degree pruning reduce later search work. The algorithm explanation makes these stages inspectable and discusses why graph construction can stop short of exact neighbors. This is a good contrast to incremental insertion-based graph builders. Documentation examples illustrate behavior; they do not establish recall guarantees for arbitrary data.

24. FALCONN-LIB/FALCONN

C++ with Python bindings — historical research/reference LSH implementation. Included for algorithmic and API design; a current maintenance cadence was not established.

  • C1: Angular collision probabilities, randomized rotations, sparse feature hashing, and query-local state make numerical assumptions and concurrency boundaries explicit. The LSH-family guide distinguishes theoretical angular behavior from applicability to other distances.
  • C2, C3: Hyperplane and cross-polytope hash families support dense/sparse data and multi-probe lookup. Pseudorandom Hadamard-based rotations reduce the cost of dense random rotations. The release history documents decoupling query objects for multithreaded querying and changes to the bit-packed table implementation. These are concrete study opportunities even without assuming contemporary production support.

25. puffinn/puffinn

C++ with Python bindings — research-oriented adaptive LSH library. Distinct from FALCONN in exposing a memory budget and requested recall/probability target as core search controls.

  • C1: Collision-probability estimation and stopping rules connect implementation behavior to probabilistic retrieval goals. The documentation distinguishes the optimized default candidate filter from a simpler mode intended for stricter expected-recall requirements; a target is not an unconditional guarantee about every configuration. Architecture and API.
  • C2, C3: Similarity, hash family, hash source, and filtering policy are separate abstractions. Independent, pooled, and tensored hash generation trade computation against hash quality. Angular and Jaccard search share the framework, with memory limits influencing query work. Inserted points become searchable after rebuild, another important lifecycle constraint. This is a research implementation, not a persistent vector database.

Exact spatial search and classical algorithm selection

26. jlblancoc/nanoflann

C++ — header-only exact kd-tree search. Although derived from FLANN, it warrants a separate entry: its data adaptors, template-based dispatch, exact-search scope, and subsequent evolution are substantive differences.

  • C2, C3: Dataset adaptors avoid copying application data into a separate matrix; compile-time dimensionality and CRTP/inlining reduce dispatch and loop overhead. Radius callbacks avoid materializing unnecessary result containers. The repository explains these design choices.
  • C1, C4: The 2011–2026 changelog records sustained changes to persistence, threading, portability, and numerical behavior. Recent entries address unsigned-coordinate arithmetic, malformed serialized indexes, reconstruction of structural invariants on load, and tests that compare distances rather than assuming a tie-break order. These make the project particularly valuable for studying exactness beyond ordinary floating-point examples.

27. flann-lib/flann

C++ with multiple bindings — classical ANN library and automatic index selection. A historical algorithm-design reference; inclusion does not imply a current release/support cadence.

  • C2: A distance-templated NNIndex interface supports alternative implementations and an autotuned index that wraps the selected method. This allows application code to use a common search interface while the library chooses an index strategy.
  • C3: AutotunedIndexParams separates target precision, build-time weight, memory weight, and sampling fraction. The autotuning implementation makes multi-objective algorithm selection concrete. The changelog additionally records radius-search bugs, portability reversions, bounds checks, and templated API changes. Read it alongside the algorithm to see why a reusable numerical library needs more than a good search heuristic.

28. KristofferC/NearestNeighbors.jl

Julia — exact kd-tree, ball-tree, and brute-force nearest-neighbor search. Adds a scientific-computing perspective and a native Julia implementation.

  • C2, C3: A shared tree interface works with Distances.jl metrics, vector or matrix inputs, batched queries, and allocation-reusing operations. Leaf size balances traversal against distance calculations, while point reordering improves locality. Metric restrictions differ between kd-trees and ball trees; the README explicitly cautions against very high-dimensional workloads.
  • C1: The KNN implementation checks invalid result counts and NaNs, handles self-exclusion and skipped candidates, translates reordered IDs back to original positions, and separates internal from returned distance types. These are concrete result and numerical invariants that are easy to overlook in a minimal tree implementation.

Coverage, search process, and limitations

Discovery used live web search with 17 formulations across these angles: distributed database architecture; disk indexes and SPANN; native C++/Rust/Go/Java libraries; GPU/CAGRA and quantized scoring; PostgreSQL and SQLite extensions; kd-trees and Julia; NN-Descent; LSH families; and newer reduced-rank search. Follow-up searches surfaced LoRANN and PUFFINN; later queries increasingly returned benchmark adapters, related experiment repositories, and implementations already covered. The selection therefore extends modestly beyond 25 entries to preserve those distinct algorithm families.

Every retained repository's GitHub page was opened, and at least one additional primary document, source file, or release/changelog was opened and read. Namespace redirects were resolved for VectorChord, Lance, NGT, and JVector; cuVS uses its verified NVIDIA location. ScaNN is one explicitly scoped monorepo entry. Nanoflann's independent implementation/evolution is explained rather than counted as an undifferentiated FLANN fork. No entry is presented as an official mirror or marked archived without evidence.

Exclusions include hosted-only offerings without an inspectable engine, thin client/wrapper repositories, tutorial RAG applications, awesome-lists, and benchmark harnesses such as VIBE. LanceDB was not separately counted because this selection focuses on the underlying Lance index/storage implementation. General-purpose search engines and broader ML monorepos were not exhaustively surveyed; this is a category selection guide, not a complete inventory. Additional promising research implementations could qualify, but are not implied to have received the same verification.

The criterion explanations distinguish source facts from the judgment that a subsystem is worth studying. No candidate code was executed, dependencies installed, or benchmark results independently reproduced. Reported performance mechanisms are supported; comparative speedups, universal recall guarantees, and uniform maintenance claims are deliberately absent. Some GitHub raw-file and documentation fetches failed transiently; retained claims use successfully read alternate primary pages. Links to default branches and unversioned documentation can change after this research date. Historical/research selections and DiskANN's unmaintained C++ branch are explicitly identified above.

Continue exploringBack to the collection →