RustFS Metadata — Final Greenfield Design (EN)
RustFS Metadata — Final Greenfield Design (EN)
Document: RUSTFS-MD-DESIGN-005 (canonical / single source of truth)
Status: DRAFT
Version: 1.0.0
Premise: No historical data. Greenfield build. No migration, ever.
Folds in: RUSTFS-MD-DESIGN-002 (core) + 004 (no-coexistence), and RETIRES the
migration/coexistence content of CHG-001 and DESIGN-003 entirely.
Language: English (中文镜像见 rustfs-final-greenfield-ZH.md)
1. New Premise and What It Removes
There is no historical data. The system is built fresh. Therefore:
REMOVED ENTIRELY (no longer part of any plan):
✗ All migration (no rustfs-migrate tool, no re-stripe, no Option A/B choice)
✗ xl.meta in every form — not at runtime, not as a migration input, not at all
✗ Every legacy-decode open item (O1 old-bitrot, O5 reproduce-old-EcDist, ...)
✗ Any backward-compatibility scaffolding
RESULT: the design is purely forward-looking. The only data the system ever sees
is data it wrote itself, in its own format.
2. Reframed Principle: "Minimal Change" → "Maximal Reuse"
The earlier goal "minimal change to the RustFS chunk" existed to avoid breaking running data. With no historical data, it is restated:
Reuse the proven, hard-to-get-right algorithms (Reed-Solomon, bitrot, disk I/O, heal reconstruction). Build a clean, formula-native chunk service on top of them. Carry over none of the object-store / xl.meta semantics.
REUSE as libraries (do not rewrite): ecstore-core (RS encode/decode),
bitrot reader/writer, disk I/O layer,
heal reconstruction math, quorum logic.
BUILD fresh (clean, no legacy): the chunk service wrapper, on-disk chunk
format, formula placement, KV metadata.
DROP forever: SetDisks object semantics, xl.meta,
object→set name hashing, DDir, EcDist
storage, versioned-object-dir layout.
Because we define the on-disk chunk format from scratch, we make it self-verifying (inline streaming bitrot) by design — so metadata never needs to carry per-block hashes, and the decode path needs only the layout params.
3. Canonical Architecture (the whole thing, one page)
CLIENT STATELESS MDS TRANSACTIONAL KV
┌──────────────────┐ ┌────────────────────┐ ┌──────────────────┐
│ metadata RPC stub │──control─▶ FS semantics │──txn─▶ INOD / DENT / │
│ placement engine │ │ inode-range alloc │ │ VIDX / VOLU / │
│ cluster-map cache │◀─layout──│ placement engine │ │ EXTM / PEND/DELQ │
│ chunk I/O ────────┼─data─┐ │ (NO data path) │ │ (TiKV / embedded)│
└──────────────────┘ │ └────────────────────┘ └──────────────────┘
▼
┌──────────────────────────────────────────┐
│ CHUNK SERVICE (fresh wrapper) │
│ · formula-native, chunk-key addressed │
│ · self-verifying on-disk chunk format │
│ reuses: RS · bitrot · disk I/O · heal │
└──────────────────────────────────────────┘
3.1 Metadata = KV rows
INOD | inode_id → InodeAttr { size, mode, uid, gid, times, kind, nlink,
body, sys_meta }
body = Inline(bytes) # ≤ ~1 KB only (rare)
| Slice{container_id,off,len} # small file → packed
| OwnLayout(Layout) # large file → own chunks
DENT | parent_ino | name → { child_ino, kind }
VIDX | bucket | key → [ { version_id, inode_id, mtime, delete_marker } ]
EXTM | volume_id | extent → ExtentEntry (block volumes — the only index)
VOLU | volume_id → Volume
CMTA | container_id → ContainerMeta { layout, capacity, used, dead_bytes,
sealed } (small-file packing)
IALOC → batched inode counter
PEND | file_id → write intent (crash GC)
DELQ | file_id → delete tombstone (chunk GC)
MPUP/MPRT | upload_id... → multipart staging
KV holds metadata only — never bulk file data. Inline is reserved for
degenerate ≤ ~1 KB cases; it is NOT the small-file strategy (see §3.5). Billions
of small files are normal for AI datasets and code/dependency trees
(node_modules, pip/conda) — their metadata (inode + dentry rows) scales in the
KV exactly as 3FS demonstrates, but their data must never go inline.
3.2 Layout = pure formula (one descriptor per file, inline in inode)
struct Layout {
file_id: u64, ec: EcParams, stripe_unit: u32, chunk_size: u64,
placement_seed: u64, cluster_map_epoch: u64,
}
// All locations computed, nothing stored per chunk:
// set = HRW(file_id, chunk_index, seed, cluster_map[epoch])
// path = deterministic(file_id, version, chunk_index)
// dist = deterministic(file_id, chunk_index)
// hashes = inline in the shard (self-verifying)
3.3 Data plane = client ↔ chunk directly (MDS never in the path)
open(): client gets Layout + cluster-map epoch once → caches (layout is
immutable for write-once files → cache for the file's lifetime)
read(): client computes location locally → ReadChunk(chunk_key, off, len)
direct to the chunk server. Zero MDS traffic.
write(): client → WriteChunk direct to chunk server (RS encode inside the
service); then MDS CommitWrite(size, mtime) — metadata only.
3.4 The clean primitives kept simple
Tiny files (≤~1KB) → inode.body = Inline (rare; degenerate optimization only).
Small files → PACKED into containers (§3.5) — the answer for billions of
small files. inode.body = Slice{container_id, off, len}.
Large files → inode.body = OwnLayout (the §3.2 formula path).
Versioning → VIDX; each version its own inode + deterministic chunk paths.
Multipart → stage parts; on Complete write formula chunks (re-stripe only
if part boundaries misalign — uncommon).
Block volumes → EXTM extent map (thin/COW), the single enumerated index.
Crash consistency → data-first, then KV commit; PEND intent + GC scrubber.
3.5 Small-file packing (first-class — billions of small files)
AI datasets (millions of tiny samples/tokens) and code/dependency trees (node_modules, pip/conda, build artifacts) routinely have hundreds of millions to billions of small files. Putting their data in the KV is fatal; sharding each tiny file with EC/replica wastes the per-shard fixed floor. The answer is packing, which composes cleanly with the formula layout:
Container = a large internal object (e.g. 256 MiB) living in a reserved internal
namespace. It IS a normal large file internally → its own OwnLayout,
formula-placed, EC'd. Reuses the large-file chunk path entirely.
Write = append the small file's bytes into the current open container for the
writer; reserve (container_id, offset, length); commit the inode with
body = Slice{container_id, offset, length}. Seal the container at its
size limit → sealed containers are immutable → ideal for EC.
Read = inode → Slice{container_id, offset, length} → container's Layout
(cached; containers are few and hot) → compute chunk location →
direct byte-range read from the chunk server. MDS off the data path.
A small slice lands within one EC data-shard fragment → EC partial
read at ~1x amplification (reuses LLD-002 random-read path).
Delete/GC = tombstone the slice; track per-container dead_bytes (CMTA). A
background compactor repacks live slices from sparse containers into
new ones and frees the old container (log-structured GC).
Scale check: 1e9 × 32 KiB files → data packs into ~131 K containers (256 MiB each) that EC efficiently (~1.5x, not the small-file shard floor); metadata is ~1e9 inode + dentry rows in the KV (~hundreds of GB, distributed — the KV's job). This is the Haystack / SeaweedFS model, adapted to our formula layout.
Costs (accepted, see PLAN-001 Track B): compaction GC; client-side batching for fast bulk small-file creation; per-container append-offset reservation (bounded by using one open container per writer/session).
4. The Chunk Service: Reuse Algorithms, Build a Clean Wrapper
One design decision to record:
DECISION: build a FRESH thin ChunkStore that calls the reused libraries directly,
rather than retrofitting the existing object-oriented SetDisks.
Why: SetDisks couples EC + distribution + xl.meta + object-level quorum/locking.
With no legacy data, retrofitting it would carry object-store assumptions
into the chunk layer. A fresh ChunkStore that composes ecstore-core (RS),
the disk layer, and the bitrot/heal modules is cleaner and formula-native.
Cost: slightly more new code than a retrofit — but it is the thorough-iteration
choice and avoids inheriting object semantics we will never use.
ChunkStore surface (chunk-key addressed, byte-range, self-verifying):
write_chunk(chunk_key, offset, data, sync) -> ()
read_chunk(chunk_key, offset, length) -> bytes (verifies inline bitrot)
delete_chunk(chunk_key) -> ()
// RS encode/decode, shard placement (by formula), bitrot, heal: reused libs
5. Build Plan (no migration — just build)
B1 rustfs-kv: TransactionalKv trait + embedded backend (redb/RocksDB) + TiKV.
B2 rustfs-inode: InodeAttr, Layout, DentValue, Volume, ExtentEntry types.
B3 rustfs-placement: shared HRW placement + cluster-map (used by MDS & client).
B4 rustfs-mds: stateless service, inode-range allocator, FS-semantics txns,
gRPC (namespace + OpenLayout + GetClusterMap + CommitWrite + vol).
B5 chunk-service: fresh ChunkStore reusing RS/bitrot/disk/heal; self-verifying
on-disk chunk format; chunk-key addressing.
B5b packing: container manager (append/seal/Slice), small-file pack/unpack,
compaction GC. Reuses B5 (containers are large internal files).
B6 rustfs-client / FUSE: placement engine, cluster-map cache, direct chunk I/O,
read-ahead (sequential) + random-read path (LLD-002);
client-side small-file write batching into container appends.
B7 S3 + block (iSCSI/NVMe-oF) front ends over the same MDS + ChunkStore.
B8 HA: point MDS at TiKV → distribution + HA for free. No custom Raft.
No phases for seam/dual-write/cutover/migrate — those concepts are gone.
6. The Few Real Open Questions Left
Q1 KV backend: TiKV (Rust-native, clustered, recommended) vs embedded
redb/RocksDB (single-node). Ship embedded first, add TiKV for HA.
Q2 inline cap (degenerate ≤~1 KB only — bulk small files go to packing, §3.5,
NOT inline). Keep inline tiny so the KV never holds bulk data.
Q3 Define the on-disk chunk format: stripe_unit, inline-bitrot block size,
chunk-key path scheme. (Free choice — no legacy constraint.)
Q4 Default storage policy by SIZE × ACCESS-MODE:
tiny ≤~1KB → Inline; small → packed containers (§3.5); large ≥~1MiB → EC
own-layout; random/partial-write (block volumes, mutable files) → replica
(factor = m+1 so all tiers tolerate the same failures).
Q5 ChunkStore: confirm the fresh-wrapper decision vs SetDisks retrofit (§4).
Q6 Container sizing + packing locality (pack by writer/temporal, optionally by
directory subtree) + compaction trigger thresholds (dead_bytes ratio).
These are forward design choices, not compatibility constraints.
7. Summary
Premise : no historical data → greenfield → no migration, no xl.meta, anywhere.
Principle : reuse proven RS/bitrot/disk/heal; build a clean formula-native shell.
Metadata : stateless MDS over a transactional KV; INOD/DENT/VIDX rows. KV holds
metadata ONLY — never bulk file data.
Layout : pure formula; zero per-chunk metadata; self-verifying shards.
Small files: PACKED into large containers (Slice pointer in the inode); billions
of small files scale; never inline bulk data in the KV.
Data plane: client ↔ chunk directly; MDS never on the data path.
Chunk : fresh thin ChunkStore reusing the hard algorithms; no object/xl.meta.
HA : TiKV provides distribution + consistency; no custom consensus.
Coexistence/migration : none. The system only ever sees its own format.
This is the end state of the simplification: a single native model, the proven math reused as libraries, and nothing carried forward.
Chinese mirror: rustfs-final-greenfield-ZH.md. Canonical design; retires the migration/coexistence content of RUSTFS-CHG-001 and RUSTFS-MD-DESIGN-003.