Category report

Object storage systems

Research date: 2026-10-09

This report selects 20 GitHub repositories implementing object/blob storage engines, substantial object-serving gateways, or a versioned object-storage layer. It covers distributed and single-node designs, replicated and erasure-coded storage, centralized and decentralized deployments, and Go, C++, Python, Java, Rust, Erlang, and JavaScript communities. Filesystem monorepos appear only where their object-storage subsystem is substantial and explicitly identified. SDKs, deployment operators, benchmark-only projects, and applications that merely consume object storage are outside scope.

These are engineering study selections, not deployment recommendations or a claim that every component is exemplary. Criterion assignments are grounded judgments based on the linked primary material. Design documents describe intended mechanisms; they are not independent proof of correctness or performance. Repository pages were opened to verify the canonical URLs, and additional primary material was read for every selection. Historical projects and official mirrors are identified below; inclusion otherwise makes no blanket claim of active maintenance.

Criteria legend

  • C1 — Difficult correctness: invariants, concurrency, adversarial inputs, consistency, durability, or failure recovery.
  • C2 — Reusable abstractions: substantial interfaces or architectural layers supporting multiple workloads, protocols, or storage backends.
  • C3 — Performance with structure: concrete responses to I/O, memory, network, throughput, or scaling constraints within an understandable architecture.
  • C4 — Sustained evolution: evidence across years of compatibility work, testing, migrations, or complexity management, rather than repository age alone.

Distributed storage engines

ceph/ceph

Language / role: Primarily C++; RADOS distributed object storage and the RADOS Gateway within the larger Ceph monorepo. Counted once, including its object engine and gateway.

Study how object placement, membership changes, and recovery fit together beneath several storage interfaces. The useful conceptual boundary is between an object's logical pool, its placement group, and the OSDs currently responsible for that group.

  • C1: Peering reconciles the state of a placement group's objects and metadata; the distinction between agreeing on state and actually possessing current data matters during recovery. Acting sets, primary ownership, and scrubbing make failure handling explicit.
  • C2 / C3: Placement groups decouple objects from physical OSDs, while clients compute placement from CRUSH and the cluster map. This avoids a per-object location service and allows topology changes without embedding fixed server locations in clients. RADOS also underlies the object, block, and file interfaces.

Read first: Dynamic cluster management explains placement, peering, and scrubbing; the architecture map situates RADOS, librados, and RGW. These are development-version documents.

openstack/swift

Language / role: Python; distributed object storage. Official GitHub mirror: the repository identifies OpenDev as the development home.

Swift is particularly instructive about eventual consistency at different layers: an object may be readable while a container listing still awaits an asynchronous update.

  • C1: Timestamped object versions and deletion tombstones prevent old replicas from resurrecting deleted objects. Proxy handoffs, replication, auditors, and erasure-code reconstruction address distinct failure cases rather than hiding them behind one generic retry loop.
  • C2: Storage policies expose different replication, hardware-tier, or erasure-code choices through separate object rings while preserving the client-facing container abstraction.
  • C3: Proxies stream object bytes without spooling; weighted rings balance capacity and constrain movement during rebalancing. The architecture explicitly discusses the tension between capacity balance and failure-domain dispersion.

Read first: Swift architectural overview, especially the ring, object server, updaters, and auditors sections.

seaweedfs/seaweedfs

Language / role: Primarily Go, with Rust components; blob storage, filer metadata, and S3 serving within a broader storage monorepo.

Study a design that keeps the master responsible for volumes rather than individual blobs. The blob identifier supplies a volume identifier; clients resolve or cache the volume-to-server mapping and access storage nodes directly.

  • C2: The master/volume blob layer can be used directly, while the filer adds names and directories over configurable metadata stores. S3 and filesystem interfaces build on these layers rather than requiring separate payload stores.
  • C3: Packing blobs into append-only volumes avoids one filesystem inode per small object. Per-volume indexes and cached location mappings reduce lookup work, while volume-level compaction and cloud tiering exploit the same layout.
  • C1: Replica placement is expressed at volume level with rack and datacenter distinctions, making failure-domain policy part of allocation rather than an incidental property of a copy loop.

Read first: Blob Store Architecture. Its request walkthrough and master/volume explanation are more useful for study than headline throughput comparisons.

apache/ozone

Language / role: Java; distributed object storage with S3 and Hadoop filesystem access. Relevant subsystems include Ozone Manager, Storage Container Manager, datanodes, and the S3 gateway.

Ozone provides a useful example of separating namespace management from block placement. Trace a key write through authorization, block allocation, direct datanode I/O, and the final metadata update.

  • C1: Ozone Manager tracks open keys, committed key metadata, multipart state, and pending deletions separately. Block tokens authorize data access, and Ratis provides the replicated metadata-state foundation. These boundaries expose the correctness work between allocating bytes and publishing an object.
  • C2: Volumes, buckets, and keys form a reusable namespace above container/block storage. The manager's service surface supports object CRUD, multipart operations, filesystem operations, and access-control management over that shared foundation.

Read first: Ozone Manager architecture, including its write/read sequence and persisted-state tables.

cubefs/cubefs

Language / role: Go; the BlobStore erasure-coding subsystem and its ClusterMgr within the CubeFS monorepo. The older standalone cubefs-blobstore repository is not counted separately.

The most distinctive study target is the integration of a consensus library with a real metadata service: disk registration, volume ownership, allocation, and background tasks all need persistence and recovery semantics.

  • C1: ClusterMgr submits state changes through Raft, persists a WAL, applies changes to RocksDB-backed state, and restores through snapshots plus subsequent log entries. Its ReadIndex path waits for sufficiently applied state after confirming leadership.
  • C3: The design batches read-index work, uses sequential WAL writes, and prunes recoverable history through persisted state and snapshots. These mechanisms address latency and restart cost without conflating the consensus engine with resource-management modules.
  • C2: Volume, disk, identifier-scope, service-discovery, and configuration managers sit behind explicit service, state-machine, protocol, persistence, and transport layers.

Read first: Raft in the CubeFS erasure-code system. Treat its broad scaling claims as project claims; the concrete WAL, ReadIndex, and recovery mechanisms are the selection evidence.

linkedin/ambry

Language / role: Java; distributed blob/object storage, originally oriented toward media objects.

Ambry is a strong place to study the relationship between an append-only blob log and the index that makes it queryable. Its local-store design is understandable independently of the full distributed system.

  • C1: The data log doubles as a recovery log. Startup scans from the last checkpoint to rebuild missing index entries; deletion records follow the original writes so replay restores deletion state as well. Corrupt Bloom filters can be rebuilt.
  • C3: Sequential writes into preallocated logs, a recent in-memory index segment, memory-mapped older segments, and Bloom filters trade memory against random I/O. Direct log-to-socket reads and streaming uploads reduce heap buffering and preserve page-cache capacity.

Read first: The Store design. It is an older design document, so use it as an architectural introduction and check implementation details against the current tree before making operational assumptions.

NVIDIA/aistore

Language / role: Go; distributed object storage for data-intensive workloads, including integration with remote object stores.

Study how deterministic placement and decentralized object movement coexist with a control-plane leader. The rebalance design is especially useful for understanding what an online read must do while ownership changes.

  • C1: A new cluster-map version changes the expected owner of some objects. During migration, the new owner can obtain an object from a neighboring target that still holds it, instead of treating incomplete movement as permanent absence.
  • C3: Rendezvous hashing computes target ownership from the cluster map, qualified bucket, and object name. Targets independently traverse and move their own objects, avoiding a central per-object transfer coordinator; separate intracluster networking can isolate movement traffic.

Read first: Global Rebalance, particularly placement, serving reads during migration, and the distinction between rebalance and local-disk resilvering.

storj/storj

Language / role: Go; satellite and storage-node components of a decentralized cloud object-storage network. Counted once; client SDKs and design-document repositories are not additional selections.

A valuable contrast to datacenter-only systems: coordination and storage nodes have separate failure histories, so local deletion policy must account for a coordinator being restored from backup.

  • C1: The piece-store implementation defaults to moving deleted pieces into trash and explicitly connects premature deletion to audit failures after a satellite database restore. Storage-directory identity verification and support for older piece formats expose further operational invariants.
  • C2 / C3: Piece data, expiration metadata, accounting, and file walking have separate interfaces. Expiration records can retain piece size to avoid another filesystem stat, while lower-priority file-walker subprocesses keep garbage collection and accounting scans from monopolizing foreground I/O.

Read first: Storage-node piece store, especially PieceExpirationDB, StoredPieceAccess, Config, and StoreForTest.

Smaller deployments and alternative implementation choices

deuxfleurs-org/garage

Language / role: Rust; S3-compatible storage for small, geographically distributed deployments. Official GitHub mirror: development is hosted on Deuxfleurs' forge.

Garage is useful for studying replicated metadata and garbage collection without assuming a conventional centralized metadata leader.

  • C1: Requests establish a metadata quorum before using object data. The internals discussion explains why CRDT tombstones cannot simply disappear: an older replica can otherwise reintroduce a deleted value. It also analyzes races between tombstone collection, partition movement, and reference-counted block deletion.
  • C3: Request routing prefers the local node, then the same zone, then lower-latency nodes. Small objects can be answered from inline metadata; larger ones require block retrieval. These are concrete optimizations for heterogeneous geographic deployments.

Read first: Internals. The garbage-collection section includes historical bug discussions and timing assumptions; it should not be read as a formal proof of the current implementation.

rustfs/rustfs

Language / role: Rust; S3-compatible distributed object storage with an erasure-coded storage engine.

Study explicit on-disk contracts alongside an implementation undergoing architectural refactoring. The project documents both intended crate boundaries and places where the current code still violates them.

  • C1: The erasure-coding contract specifies shard geometry, checksums, read/write quorum, healing, and metadata-selected decoding. It distinguishes modern and legacy codec formats, making backward decoding a per-object concern rather than a process-wide switch.
  • C2: The request path separates HTTP serving, use-case orchestration, storage adaptation, the erasure engine, and I/O support. Dedicated crates expose security, policy, metadata, and storage concepts beyond a single request handler.
  • C3: The documented I/O layer includes buffer pools, admission control, and backpressure. Its known structural issues are useful study material rather than evidence of uniformly clean layering.

Read first: Architecture and erasure-coding contract. No C4 claim is made; compatibility intent does not by itself establish years of successful evolution.

uroni/hs5

Language / role: C++; single-node S3-compatible object storage, deliberately designed to scale up.

HS5 offers a compact counterpoint to distributed engines: an LMDB index maps object names to offsets in a shared data file, with free-space management in metadata.

  • C1: The documented publication sequence syncs object data before syncing its index entry. Default acknowledged-write durability is explicitly distinguished from manual-commit mode. Release 1.0.0 records fixes for races between reads and deletion, including multipart reads.
  • C3: Append-heavy writes use the shared data file sequentially, while listing and most deletion bookkeeping use the metadata database. Index and payload files can live on different media, making the small-object and metadata-I/O tradeoff visible.

Read first: The repository's storage and durability explanation and release notes. This is a single-node design; the report does not infer distributed redundancy. Release notes are also worth checking against README feature omissions, which can lag implementation.

Object-serving gateways and storage composition

scality/cloudserver

Language / role: JavaScript / Node.js; Zenko CloudServer, an S3 server with multiple data-backend options. This is not the proprietary Scality RING implementation.

Study how protocol-level versioning semantics can be layered over a simpler metadata CRUD service.

  • C1: Each version has a metadata key, while a special current-version entry accelerates ordinary GET. The design distinguishes null versions, delete markers, enabled/suspended versioning, and updates arriving through replication; those states change how the current-version entry must be maintained.
  • C2: S3-specific policy lives in the API layer, while the metadata backend accepts explicit CRUD instructions for versions. The separate data-metadata daemon exposes metadata RPC and data REST interfaces, allowing several connectors to share stored state.
  • C3: The current-version entry avoids scanning all versions on ordinary reads, and the daemon batches metadata updates before database writes.

Read first: Architecture and versioning design. The cited document is on the verified development/9.5 branch.

noobaa/noobaa-core

Language / role: JavaScript with native components; S3 data service spanning local filesystems and cloud object stores.

NooBaa is useful for studying object-storage composition: a logical bucket can represent storage behind different providers, while placement and replication remain user-visible policies.

  • C2: Data buckets and namespace buckets share replication-policy concepts, including destination buckets and key-prefix filters. Bucket classes can supply policies to subsequently created bucket claims, separating reusable policy from individual storage resources.
  • C3: The documented baseline replication mechanism compares source and destination listings. A log-assisted path prioritizes recent changes while the broader scan catches up, exposing the tradeoff between complete reconciliation and low-latency incremental work.
  • C1: Deletion and version propagation have explicit provider- and mode-dependent semantics; they cannot safely be inferred from the word “replication.” The current guide scopes these options to supported AWS log-based configurations.

Read first: Bucket replication guide and original replication design. The older design excludes versions and deletions; the newer guide documents later, limited support.

versity/versitygw

Language / role: Go; S3 gateway over filesystems and other storage backends.

This is a particularly concrete codebase for studying the impedance mismatch between object semantics and a transparent POSIX namespace.

  • C1: Uploads write temporary files and publish completed values into the namespace, avoiding partially overwritten object contents. Conditional writes must serialize the final precondition check with publication. The documentation distinguishes cross-process filesystem locks from in-process locks and states the filesystem coherence assumptions.
  • C2: S3 operations are translated through backend implementations, with POSIX metadata stored through xattrs or a sidecar layout. This supports use as a storage service over existing infrastructure rather than requiring a new payload engine.
  • C3: Pooled I/O buffers, concurrency limits, kernel-assisted multipart copying, and optional direct I/O make throughput and page-cache tradeoffs explicit.

Read first: POSIX Backend, especially conditional publication, temporary uploads, metadata, and direct I/O. Its versioning feature is explicitly marked experimental in the inspected documentation.

gaul/s3proxy

Language / role: Java; embeddable S3 server and translation gateway to multiple storage backends.

S3Proxy is more than a forwarding wrapper: it implements S3 request semantics when an underlying provider exposes different operations and error behavior.

  • C1: Multipart completion validates checksum headers, looks up uploaded parts, checks part numbers and available ETags, and handles completion retries and missing upload state. The implementation contains explicit Azure/GCS paths and documents cases where backend behavior limits validation or atomicity.
  • C2: An HTTP-server-independent handler operates against a BlobStore abstraction. The project supports embedding, middleware extensions, filesystem/in-memory uses, and several remote providers, making protocol logic reusable across deployment shapes.

Read first: S3ProxyHandler.java, particularly handleCompleteMultipartUpload and provider-specific handling. Do not assume perfect semantic equivalence across backends: the source explicitly calls out an Azure conditional-write emulation using a non-atomic HEAD/PUT sequence.

Versioned object-storage layer

treeverse/lakeFS

Language / role: Go; a versioned namespace and S3-facing layer over existing object storage, rather than a replacement disk or replication engine.

Study how immutable object payloads, immutable committed metadata, and mutable branch references combine to support branching and merging over large object collections.

  • C1: Logical names resolve through staging state or a committed snapshot; overwriting an object creates a new physical address. Mutable references need strong consistency, while three-way merge distinguishes unchanged, changed, and deleted objects and detects conflicting outcomes.
  • C2 / C3: Graveler represents committed metadata as content-addressed SSTable ranges and meta-ranges. Unchanged ranges can be reused across commits, while immutable caching avoids complex invalidation. A generic KV interface separates mutable metadata operations from the backing database.

Read first: Internals, especially versioning, ranges/meta-ranges, and the metadata-store sections. The page mixes shared architecture with explicitly labeled Enterprise capabilities; this selection does not attribute Enterprise-only async merge or layout features to the open-source repository.

Historical designs and specialist service stacks

minio/minio

Language / role: Go; S3-compatible object storage. Archived: GitHub states that the owner archived the repository on 2026-04-25. Included as a substantial historical implementation, not as an actively maintained community server.

MinIO is valuable for studying object-level erasure coding and the consequences of making erasure sets the unit of quorum and repair.

  • C1: Read/write quorum and healing are scoped to the object's erasure set. Drive ordering and host distribution affect failure tolerance; pool expansion must preserve the shared namespace and avoid conflicting object creation.
  • C3: The design deliberately bounds erasure-set size to limit communication overhead. Objects map to sets by hashing, while new-object placement across server pools considers utilization. Per-object storage classes expose a redundancy/capacity tradeoff without changing the whole volume layout.

Read first: Distributed Server Design Guide. Read its geometry and expansion rules in the context of the archived code version.

leo-project/leofs

Language / role: Erlang; distributed, eventually consistent object storage. This entry studies the historical v1 system and its component integration; it does not establish current maintenance status.

LeoFS separates gateway caching and protocol handling, object storage with replication/recovery queues, and manager supervision of nodes and ring state.

  • C1: The changelog records concrete correctness work around ring inconsistency, rack-aware replication, multipart-abort cleanup, data-sync error handling, and large-object recovery. These identify realistic failure paths to trace rather than merely advertised availability.
  • C4: The inspected changelog spans releases from 2014 through 2019, including state migration fixes, S3 client interoperability changes, and adaptation to newer Erlang/OTP versions. This is evidence of sustained compatibility and operational repair work during that period.
  • C2: The gateway/storage/manager split and reusable Leo libraries expose separate caching, routing, object-storage, and queueing responsibilities.

Read first: The repository architecture overview and v1 changelog. Several external legacy documentation links failed to load during research, so the retained evidence comes from GitHub.

basho/riak_cs

Language / role: Erlang; historical S3-compatible cloud storage built over Riak KV, also branded Riak S2 in the inspected release notes. Current maintenance is not established here.

Study the distinction between an available object/block data path and the special coordination needed for globally unique buckets and users.

  • C1: Large objects are split into blocks described by manifests. The official architecture explanation identifies a request serializer for globally unique entities, with a different failure boundary from ordinary object requests. Release notes describe races between garbage collection and full-sync replication that could resurrect deleted blocks.
  • C3: Garbage-collection refinements include bypassing the background collector for small-object deletion, adjustable worker concurrency, and explicit collection windows. These changes expose how object-size distribution and replication activity affect maintenance throughput.
  • C2: Riak CS builds object manifests, S3 tenancy, and bucket semantics over an existing distributed KV layer rather than implementing every persistence mechanism inside the protocol server.

Read first: The official architecture comparison, using its Riak-specific descriptions, and the 2.1 release notes. The latter explicitly documents compatibility boundaries and GC changes.

TritonDataCenter/manta-muskie

Language / role: JavaScript / Node.js; the Manta WebAPI, specifically the Directory API front door in the wider Manta service stack. The documentation-only Manta umbrella repository is not counted as another implementation.

This offers an object-storage API design outside the S3 lineage: whole-object HTTP operations, directory semantics, signed URLs, and multipart uploads, integrated with a separate metadata and storage tier.

  • C1: Manta's architecture separates strongly consistent namespace metadata from replicated whole-object data and explicitly chooses consistency over availability during network partitions. The API layer must preserve these boundaries while handling checksums, authentication, and interrupted uploads.
  • C3: Muskie documents request throttling through a bounded concurrency queue, probes for queue latency and client disconnects, and avoidance of high-cardinality metric labels that would exhaust memory. These are unusually concrete examples of performance observability inside an object gateway.
  • C2: The front door is separated from sharded metadata services and storage servers, allowing those tiers to scale and evolve independently. It requires the wider Manta stack; it is not a standalone S3 server.

Read first: The REST API and the official Manta architecture/design overview. The repository README supplies the throttling and metrics details. Some documents retain older Joyent terminology, and their existence is not evidence of current service availability or maintenance cadence.

Search coverage and limitations

Discovery used well over six distinct formulations, including distributed S3/erasure-coding engines; Ceph/Swift/Ozone architecture; Rust and geographically distributed stores; Erlang stores and Riak manifests; JVM media/blob systems; C++ scale-up stores; CubeFS BlobStore and AI-oriented storage; decentralized object storage and audit/repair; POSIX-to-S3 gateways; multicloud namespace/replication systems; and versioned object namespaces. Later alternative-oriented searches largely repeated these families or surfaced wrappers and closely related forks, giving diminishing returns. Search results were used for discovery; the retained technical claims link to opened primary sources.

Important scope decisions:

  • OpenIO SDS: investigated through official documentation and GitHub discovery, but the documented open-io/oio-sds repository returned 404. No substantive official GitHub mirror was verified, so it is excluded. This is a verification limitation, not a judgment about the original architecture.
  • No duplicate implementations: Ceph, CubeFS, SeaweedFS, and Storj are counted once each. Separate gateways, operators, SDKs, and old component repositories from those families were not used to inflate the count. Garage wrappers and MinIO-derived community forks were not retained without separately establishing substantive independent evolution.
  • Boundary cases: gateways remain in scope because they implement meaningful object-service semantics. lakeFS is labeled as a versioning layer. Generic databases, object-store-backed analytics engines, filesystem clients, tutorials, and curated lists were excluded.
  • Evidence limits: this was read-only source/documentation research, with no cloning, dependency installation, code execution, benchmarks, or fault injection. Some GitHub code views and older external documentation failed to render; entries were retained only when other primary material supplied the required evidence. Historical design notes are labeled where their age or assumptions matter. No unsupported throughput rankings, durability percentages, or blanket S3-conformance claims are made.
  • Evolution claims: C4 is used only where a multi-year change record was actually inspected. Other repositories may also satisfy it, but a recent push, popularity, or a compatibility feature alone was not treated as proof.

The result is a selection guide to contrasting engineering decisions: placement and repair, log/index recovery, durable publication, metadata consistency, protocol translation, and reusable versioned namespaces.

Continue exploringBack to the collection →