RustFS Metadata Structure Specification (EN)
RustFS Metadata Structure Specification (EN)
Document: RUSTFS-SPEC-META-001
Status: DRAFT — FREEZE TARGET for milestone M0/M1
Version: struct_version = 1
Scope: The COMPLETE, normative definition of every metadata record: KV key
encodings and value structures, namespace organization, referential
invariants, transaction atomicity. Metadata counterpart to
SPEC-CHUNK-001. Consolidates and supersedes the scattered metadata
definitions in MD-DESIGN-002/004/005.
Language: English (中文镜像见 rustfs-metadata-structures-ZH.md)
1. Conventions
· Storage: a transactional KV (TiKV / FoundationDB / embedded redb/RocksDB).
· Keys: byte strings. First byte = TABLE TAG (§2). Multi-byte ints in keys are
big-endian (BE) EXCEPT inode/file/container ids in point-lookup tables,
which are little-endian (LE) to spread sequential ids across KV shards.
· Values: 1 byte `struct_version` (=1) then MessagePack of the struct. Readers
MUST reject an unsupported struct_version. (XATR values are opaque.)
· Strings: UTF-8. Path components MUST NOT contain '\0' or '/'.
· Time: Timespec { secs: i64, nanos: u32 } (UTC).
· Bytes: length-delimited binary.
2. KV Key Space
Tag Table Key layout (after the 1-byte tag) Value (§)
──── ────── ────────────────────────────────────────── ─────────
0x01 INOD inode_id : u64 LE InodeAttr §5
0x02 DENT parent_ino : u64 BE | name DentValue §6
0x03 VIDX parent_ino : u64 BE | name VersionList §7
0x04 XATR inode_id : u64 LE | xattr_name Bytes (opaque) §8
0x05 EXTM volume_id : u128 BE | extent_index : u64 BE ExtentEntry §10.3
0x06 VOLU volume_id : u128 LE Volume §10.1
0x07 SNAP snapshot_id : u128 LE Snapshot §10.2
0x08 CMAP epoch : u64 BE ClusterMap §11
0x09 CMAPC (singleton) u64 epoch §11
0x0A IALOC (singleton) u64 counter §12
0x0B PEND file_id : u64 LE WriteIntent §13.1
0x0C DELQ file_id : u64 LE DeleteTomb §13.2
0x0D CMTA container_id : u64 LE ContainerMeta §14
0x0E MPUP upload_id : u128 LE MultipartState §15
0x0F MPRT upload_id : u128 BE | part_no : u32 BE PartInfo §15
0x10 BUCK bucket_ino : u64 LE BucketConfig §9
0x11 STAT inode_id : u64 LE DirStat §16
0x12 QUOT inode_id : u64 LE Quota §16
0x13 IDEN principal_id : utf8 Identity §17
0x14 AKEY access_key_id : utf8 AccessKey §17
0x15 LOCK inode_id : u64 LE | start : u64 BE | len:u64 BE LockEntry §18
DENT/VIDX use BE parent_ino as a grouping prefix → a directory's children
form a contiguous range (readdir = one ordered scan; lookup = one get). LE ids in
point-lookup tables spread sequential allocation across shards (no write hotspot).
3. Identity Types & Reserved Inodes
type InodeId = u64; type FileId = u64; type ContainerId = u64;
type VolumeId = u128; type SnapshotId = u128; type VersionId = u128; type Epoch = u64;
const NIL_INODE: InodeId = 0;
const ROOT_INODE: InodeId = 1; // "/"
const FS_ROOT: InodeId = 2; // "/fs" POSIX namespace
const BUCKETS_ROOT: InodeId = 3; // "/buckets" S3 buckets
const VOLUMES_ROOT: InodeId = 4; // "/volumes" block volumes
const SYS_ROOT: InodeId = 5; // "/.sys" reserved internal
// inodes 1..15 reserved & pre-created at format; user/container ids start at 16.
file inode ids and container ids draw from the SAME IALOC counter (§12) so they never collide.
4. Namespace Organization (file + object + block in ONE tree)
Single global tree (the whole point of "unified"). Reserved layout:
/ ROOT_INODE
/fs POSIX namespace root
/buckets each child Directory = one S3 bucket (its config in BUCK, §9)
/volumes each child = one BlockVolume inode (its data in VOLU/EXTM, §10)
/.sys reserved internal (containers have NO dentry; they live only as
CMTA + formula chunks, §14)
S3 key ↔ directory mapping (UNIFIED):
Object key "a/b/c.txt" in bucket B → path /buckets/B/a/b/c.txt.
Because file & object share ONE tree, PUT of "a/b/c.txt" creates intermediate
Directory inodes a/ and b/ (mkdir -p) so a POSIX mount sees them, and
ListObjects(prefix, delimiter='/') maps to readdir of the matching directory.
A key ending '/' (directory marker) → a Directory inode with empty body.
Cost: implicit-dir creation on first PUT under a new prefix (one mkdir per new
component, in the same create transaction). Benefit: O(1)/dir prefix listings.
5. INOD — InodeAttr (central record)
struct InodeAttr {
ino: InodeId,
kind: InodeKind, // §5.1
size: u64, // logical bytes (dirs: 0)
mode: u16, // perm bits rwxrwxrwx + suid|sgid|sticky (12 bits)
uid: u32,
gid: u32,
nlink: u32, // files: # hard links; dirs: 2 + #subdirs (§20)
atime: Timespec, mtime: Timespec, ctime: Timespec, btime: Timespec,
body: InodeBody, // §5.2 — how/where data lives
vol: Option<VolumeId>, // iff kind==BlockVolume → live VOLU(vol)
sys: SysMeta, // §5.6
gen: u64, // bumped on EVERY attr/body change (cache validation)
}
5.1 InodeKind
enum InodeKind { RegularFile=0, Directory=1, Symlink=2, BlockVolume=3 }
5.2 InodeBody
enum InodeBody {
Empty, // 0 bytes / dir / block-volume
Inline(Bytes), // ≤ inline_cap (§5.3); tiny files AND symlink targets
Slice(SliceRef), // small file PACKED in a container (§5.4)
Layout(Layout), // large file with its own formula-placed chunks (§5.5)
Segmented(SegmentList), // multipart-completed object: a part table (§23.1)
}
struct SegmentList { segments: Vec<Segment> } // tiles [0,size) contiguous & disjoint
struct Segment { offset: u64, length: u64, part_file_id: FileId, layout: Layout }
Segmented is used ONLY for multipart-completed objects (large & few); a read at
object offset O binary-searches segments, then resolves within that segment's
part_file_id by formula. Single-part uploads collapse to Layout directly.
5.3 Inline cap (degenerate only)
Inline is for ≤ ~1 KiB ONLY (symlink targets, near-empty files). NOT the
small-file strategy (bulk small files → Slice/packing). Keeping inline tiny
guarantees the KV never holds bulk data. Per-backend config; default 1 KiB.
5.4 SliceRef (packed small file → container needle)
struct SliceRef { container_id: ContainerId, offset: u64, length: u32, cookie: u32 }
// cookie MUST equal the NeedleHeader.cookie (SPEC-CHUNK-001 §11).
5.5 Layout (large file's own chunks — pure formula) + Policy
struct Layout {
file_id: FileId, // == ino
policy: Policy,
stripe_unit: u32, // EC random-read amplification control
chunk_size: u64,
placement_seed: u64,
cluster_map_epoch: Epoch,
}
enum Policy {
ErasureCoded { algo: u8, k: u16, m: u16, ec_block_size: u64, csum_algo: u8 },
Replicated { factor: u8, csum_algo: u8 }, // factor = m+1 for parity
}
All chunk locations are COMPUTED from (file_id, chunk_index, version) + seed +
CMAP[epoch] (SPEC-CHUNK-001 §6/§7); nothing stored per chunk. Immutable for
write-once files → client caches it for the file's lifetime.
5.6 SysMeta (S3 / tiering / replication / encryption)
struct SysMeta {
content_type: Option<String>, content_encoding: Option<String>,
etag: Option<String>, storage_class: Option<String>,
user_meta: BTreeMap<String,String>, // x-amz-meta-* (small; large → XATR)
tags: BTreeMap<String,String>, // S3 object tags
checksums: BTreeMap<u8,Bytes>, // algo→digest (CRC32C/SHA256/…)
encryption: Option<EncryptionInfo>, // §5.7
transition: Option<TransitionInfo>, restore: Option<RestoreInfo>,
repl_status: Option<u8>, // 0 pending,1 done,2 failed,3 replica
acl: Option<String>, // object ACL (if used)
}
struct TransitionInfo { tier: String, status: u8, transitioned_version: Option<VersionId> }
struct RestoreInfo { ongoing: bool, expiry: Option<Timespec> }
5.7 EncryptionInfo (SSE)
struct EncryptionInfo {
algo: u8, // 0 none,1 AES-256-GCM,2 AES-256-CTR,3 SSE-KMS
key_ref: Option<String>, // KMS key id / keyring ref (NEVER the key itself)
iv: Bytes, // per-object nonce/IV
dek_wrapped: Option<Bytes>,// envelope: wrapped data-encryption key
}
// present ⇔ the chunks' SPEC-CHUNK-001 flags.encrypted is set; bitrot is over
// ciphertext (post-transform).
6. DENT — DentValue
struct DentValue { child_ino: InodeId, kind_tag: u8 } // kind_tag = InodeKind disc.
Key 0x02 | parent_ino(BE) | name. readdir = range scan over 0x02|parent;
lookup = one get.
7. VIDX — VersionList (S3 versioning; optional)
DENT.child_ino always points to the CURRENT version's inode (fast path); VIDX
records the full chain (only when versioning is on and history exists).
struct VersionList { latest: u32, entries: Vec<VersionEntry> } // newest-first
struct VersionEntry { version_id: VersionId, inode_id: InodeId, mtime: Timespec,
flags: u8 /*bit0 delete_marker, bit1 is_latest*/ }
8. XATR — extended attributes
Key = 0x04 | inode_id(LE) | xattr_name Value = raw bytes (opaque)
Small S3 user-meta lives inline in SysMeta; arbitrary/large POSIX xattrs use XATR.
listxattr(ino) = range scan over 0x04 | inode_id.
9. BUCK — BucketConfig
// Key = 0x10 | bucket_ino(LE) (the /buckets/<name> Directory inode)
struct BucketConfig {
bucket_ino: InodeId, name: String, owner: String, created: Timespec,
versioning: u8, // 0 disabled,1 enabled,2 suspended
object_lock: Option<ObjectLock>,
lifecycle: Vec<LifecycleRule>,
policy: Option<String>, // JSON bucket policy
cors: Option<String>,
default_enc: Option<EncryptionInfo>, // SSE default for new objects
tags: BTreeMap<String,String>,
quota_bytes: Option<u64>,
}
struct ObjectLock { mode: u8, retain_days: u32, legal_hold_default: bool }
struct LifecycleRule { id: String, prefix: String, expire_days: Option<u32>,
transition_days: Option<u32>, transition_tier: Option<String>,
noncurrent_expire_days: Option<u32> }
A bucket IS a Directory inode under /buckets; BUCK holds its S3-specific config.
10. Block Volumes
10.1 VOLU — Volume
struct Volume {
id: VolumeId, name: String, ino: InodeId, capacity: u64,
block_size: u32, chunk_size: u64, provisioning: u8 /*0 thin,1 thick*/,
state: u8 /*0 avail,1 attached,2 snapshotting,3 deleting*/,
replica_factor: u8, csum_algo: u8, created: Timespec, snapshots: Vec<SnapshotId>,
}
Volume extents are formula-placed: chunk_key = f(volume_id, extent_index, version).
10.2 SNAP — Snapshot
struct Snapshot { id: SnapshotId, volume_id: VolumeId, created: Timespec,
capacity: u64, cow_version: u64 }
10.3 EXTM — ExtentEntry (sparse extent map — the only enumerated index)
// Key = 0x05 | volume_id(BE) | extent_index(BE)
struct ExtentEntry { version: u64 /*COW→selects chunk_key*/, flags: u8 /*0 dirty,1 cow_shared*/ }
Absent extent ⇒ unwritten ⇒ reads zeros (thin). A write to a cow_shared extent
bumps version → new chunk_key → write → update entry.
11. CMAP / CMAPC — Cluster Map
struct ClusterMap { epoch: Epoch, sets: Vec<ErasureSet> }
struct ErasureSet { id: u32, weight: f64, failure_domain: u32, members: Vec<DiskRef> }
struct DiskRef { node_id: u32, disk_id: u32, endpoint: String, status: u8 } // 0 up,1 healing,2 down
CMAPC (singleton) = current epoch. Changes only on topology change. Tiny; cached
by clients & MDS (shared rustfs-placement).
12. IALOC — id allocator
Key 0x0A (singleton) → u64 `next`. Each stateless MDS reserves a RANGE of `batch`
(e.g. 16384) ids via one CAS, then hands them out from memory. file inode ids AND
container ids draw from this single counter (no collision). ids start at 16.
13. GC rows
13.1 PEND — WriteIntent
// Key = 0x0B | file_id(LE)
struct WriteIntent { layout: Layout, started: Timespec }
Written BEFORE the first chunk of a new OwnLayout file; deleted on commit_write. A scrubber deletes formula-located chunks of any stale PEND whose inode never committed.
13.2 DELQ — DeleteTombstone
// Key = 0x0C | file_id(LE)
struct DeleteTombstone { layout: Layout, size: u64, deleted: Timespec }
Written by unlink when an OwnLayout file's nlink→0. GC computes
ceil(size/chunk_size) chunks and deletes each by formula, then removes DELQ.
Packed (Slice) files need NO DELQ — unlink bumps CMTA.dead_bytes; compaction
reclaims (§14).
14. CMTA — ContainerMeta (small-file packing)
// Key = 0x0D | container_id(LE)
struct ContainerMeta {
container_id: ContainerId, layout: Layout, capacity: u64, used: u64,
dead_bytes: u64, sealed: bool, created: Timespec,
}
A container is an internal large file (formula-placed chunks, SPEC-CHUNK-001 §11), with NO inode and NO dentry. Append/seal/delete/compaction per SPEC-CHUNK-001 §11.5.
15. MPUP / MPRT — Multipart (S3)
// Key = 0x0E | upload_id(LE)
struct MultipartState { bucket_ino: InodeId, key: String, started: Timespec,
init_meta: SysMeta, policy: Policy /*intended object policy*/ }
// Key = 0x0F | upload_id(BE) | part_no(BE)
struct PartInfo { size: u64, etag: String, checksums: BTreeMap<u8,Bytes>,
part_file_id: FileId, layout: Layout } // staged self-contained
Each part is staged self-contained under its own part_file_id (UploadPart, §23.1);
final offsets are unknown until Complete. Complete builds a Segmented body from
the parts (no re-stripe) — full state machine in §23.1. Abort/expire GC all staged
part_file_ids (§23.2).
16. Accounting & Quota
// STAT — recursive subtree accounting (advisory)
// Key = 0x11 | inode_id(LE)
struct DirStat { files: u64, dirs: u64, bytes: u64, updated: Timespec }
// QUOT — quota on a subtree (bucket or directory)
// Key = 0x12 | inode_id(LE)
struct Quota { max_bytes: Option<u64>, max_inodes: Option<u64> }
DirStat is maintained lazily (recursive stats are costly to keep strictly consistent): a background aggregator or eventually-consistent counters. S3 bucket usage reads the bucket inode's DirStat. Quota checks may use the (slightly stale) DirStat — advisory, never gating correctness (I13).
17. Identity & Access (minimal; full IAM layered above)
// IDEN — Key = 0x13 | principal_id(utf8)
struct Identity { principal_id: String, display: String, created: Timespec, groups: Vec<String> }
// AKEY — Key = 0x14 | access_key_id(utf8)
struct AccessKey { access_key_id: String, secret_hash: Bytes, principal_id: String,
created: Timespec, disabled: bool }
S3/STS auth: AKEY→principal, then authorize via BUCK.policy + object SysMeta.acl. Roles/federation/inline-policy are out of v1 core scope.
18. LOCK — POSIX advisory locks (v1-optional)
// Key = 0x15 | inode_id(LE) | start(u64 BE) | len(u64 BE)
struct LockEntry { owner: u64 /*client/session*/, kind: u8 /*0 read,1 write*/, acquired: Timespec }
Transactional; conflicts resolved by the MDS. Deferred by default (AI training rarely needs byte-range locks); the record is defined so the schema is complete.
19. Mutable-file, Sparse & Truncate Semantics
OwnLayout files are NOT assumed append-only:
· SPARSE : a chunk_index never written does NOT exist; reads return zeros up to
`size`. Existing chunks ⊆ [0, ceil(size/chunk_size)).
· GROW/APPEND : writing past EOF raises size; new chunks by formula; commit_write
updates size/mtime/gen; Layout unchanged.
· OVERWRITE in place : overwrites the target chunk (EC ⇒ RMW at bitrot_block
granularity → mutable/random-write files SHOULD use Replicated policy,
DESIGN-005 Q4).
· TRUNCATE down to N : size=N; chunks fully beyond N deleted by formula; the
boundary chunk's valid tail is bounded by chunk_logical_len in the
shard header (SPEC-CHUNK-001 §3) — bytes past it read as zero.
· TRUNCATE up : size=N, no chunks written → the gap is sparse (zeros).
· `version` in the chunk path : OBJECT version for S3-versioned writes and COW; a
plain POSIX in-place write keeps the same version and overwrites. A new
S3 version creates a NEW inode (own file_id) → its own chunk paths, so
the prior version's data is untouched.
20. Referential Integrity & Invariants
I1 Every DENT.child_ino / VIDX.entries[].inode_id references a live INOD.
I2 files: nlink == #DENT pointing at it. dirs: nlink == 2 + #child dirs.
I3 Directory ⇒ body==Empty; Symlink ⇒ body==Inline(target);
BlockVolume ⇒ body==Empty ∧ vol==Some(v) with live VOLU(v).
I4 body==Slice(s) ⇒ live CMTA(s.container_id); read cross-checks s.cookie==
NeedleHeader.cookie ∧ NeedleHeader.file_id==ino.
I5 body==Layout(L) ⇒ L.file_id==ino; chunks exist by formula (or PEND pending).
I6 VIDX(p,n) exists ⇒ versioning on ∧ current entry.inode == DENT(p,n).child_ino.
I7 EXTM sparse: a missing extent reads as zeros.
I8 container_id and file inode ids never collide (shared IALOC).
I9 gen increases on every attr/body mutation.
I10 S3 object /buckets/B/<key> is RegularFile (or Directory for key ending '/');
intermediate components are Directory inodes (implicit mkdir -p).
I11 BUCK(bucket_ino) exists ⇔ bucket_ino is a direct child Directory of /buckets.
I12 OwnLayout may be sparse: existing chunks ⊆ [0, ceil(size/chunk_size)).
I13 DirStat/Quota are advisory and may lag; they gate POLICY, never correctness.
I14 SysMeta.encryption present ⇔ the object's chunks' flags.encrypted set.
I15 No directory is an ancestor of itself; rename MUST reject moving a dir into
its own subtree (loop prevention).
I16 rmdir/dir-rename-overwrite require the target directory be empty (no DENT
rows under its prefix).
I17 An inode with nlink==0 but open_count>0 is ORPHANED: unlinked from the tree,
kept live until the last close (no DENT; chunk GC deferred to final close).
I18 Hardlinks are only to RegularFile inodes within the same namespace; never to
Directory (loop) and never across the file/object/volume reserved roots.
I19 readdir returns each live entry exactly once for a cursor scan with no
concurrent rename of already-returned names; entries are name-ordered.
I20 A multipart-Completed object has body=Segmented (or single Layout if 1 part);
segments tile [0,size) contiguously & disjointly; each part_file_id's chunks
exist by formula.
I21 A delete marker (VIDX entry, delete_marker=1) ⇒ no DENT for that name; the key
is absent from POSIX / ListObjectsV2 but present in ListObjectVersions.
I22 Aborting/expiring a multipart upload removes MPUP + all its MPRT and GCs every
staged part_file_id (no orphaned staged chunks).
21. Transaction Atomicity Groups
Each is exactly ONE serializable KV transaction; data is durable on chunk servers BEFORE any metadata-commit transaction.
create(parent,name,attr) { put INOD; put DENT; bump parent.mtime,gen }
mkdir(parent,name) { put INOD(dir,nlink=2); put DENT; parent.nlink++;
bump parent.mtime,gen }
hardlink(ino,parent,name) { ino.nlink++ , ino.ctime,gen; put DENT }
symlink(parent,name,target) { put INOD(Symlink, body=Inline(target)); put DENT }
unlink_ownlayout(parent,name) { del DENT; ino.nlink-- ; if 0 { del INOD; put DELQ } }
unlink_packed(parent,name) { del DENT; del INOD; CMTA.dead_bytes += hdr+len }
rmdir(parent,name) { (empty) del DENT; del INOD; parent.nlink-- ; mtime,gen }
rename(sp,sn,dp,dn) { del DENT(sp,sn); put DENT(dp,dn); dst-overwrite +
dir-move nlink fixups; ino.ctime,gen; sp/dp mtime }
create_packed(parent,name,sr) { CMTA.used += hdr+len (reserve, writer's open
container); put INOD(body=Slice); put DENT }
// AFTER the needle is durable in the container
commit_write(ino,size,mtime) { ino.size=size; mtime; gen++ ; del PEND(ino) }
// AFTER chunks durable
truncate(ino,N) { ino.size=N; mtime,gen++ ; enqueue chunk trim > N }
setattr(ino, …) { mutate fields; gen++ }
put_object_s3(B,key,attr) { mkdir -p implicit dirs (each put INOD+DENT,
parent nlink/mtime); put object INOD+DENT;
if versioning upsert VIDX } // namespace flip in 1 txn
create_bucket(name,owner) { put Directory INOD under /buckets; put DENT; put
BUCK; BUCKETS_ROOT.nlink++ }
delete_bucket(name) { (empty) del BUCK; del DENT; del INOD; BUCKETS_ROOT.nlink-- }
SSI conflicts are retried by the MDS. Cross-directory rename is one global txn.
22. Namespace Operation Semantics (precise behaviour)
This section pins the edge behaviour the §21 sketches abbreviate.
22.1 Name rules
· A name is 1..=255 bytes UTF-8, no '\0', no '/'. "." and ".." are reserved and
never stored as DENT rows (resolved structurally).
· Case-SENSITIVE (POSIX). S3 keys are byte-exact; the same bytes map to the same
DENT. Normalization is NOT performed.
· Total path length is unbounded by the schema; clients MAY impose PATH_MAX.
22.2 lookup / resolve
resolve(path): split on '/'; from the appropriate root (ROOT/FS_ROOT/BUCKETS_ROOT),
for each component c: get DENT(cur, c) → child_ino; cur = child_ino.
"." → cur; ".." → parent (tracked by the walker; ROOT's parent is ROOT).
A symlink component: read body=Inline(target); resolve target (bounded to
SYMLINK_MAX = 40 hops; exceeding → ELOOP).
Missing component → ENOENT. A non-final component that is not a Directory →
ENOTDIR.
22.3 rename(src_parent, src_name → dst_parent, dst_name) — full semantics
Preconditions / errors (checked inside the single txn):
· src must exist (ENOENT).
· If src is a Directory and dst_parent is src or a descendant of src → EINVAL
(loop; invariant I15). The descendant check walks dst_parent upward to ROOT
looking for src's inode; bounded by tree depth.
· Sticky-bit (mode & 01000) on a parent restricts rename/delete to the entry
owner or dir owner (POSIX) — enforced if POSIX perms are on.
Overwrite of an existing dst:
· dst is a RegularFile/Symlink: it is atomically REPLACED. The old dst inode's
nlink-- ; if 0 → del INOD + (DELQ if OwnLayout | CMTA.dead_bytes if Slice).
· dst is a Directory: allowed ONLY if dst is empty (I16) AND src is also a
Directory → EISDIR/ENOTDIR/ENOTEMPTY otherwise. Empty dst dir is removed
(dst_parent.nlink--).
· src kind ≠ dst kind in the file↔dir sense → ENOTDIR / EISDIR per POSIX.
Effects (ONE txn):
del DENT(src_parent,src_name); put DENT(dst_parent,dst_name)=src_child;
if moving a Directory across parents: src_parent.nlink-- , dst_parent.nlink++
(the moved dir's '..' re-points; the dir inode's nlink is unchanged);
src_child.ctime,gen++ ; src_parent.mtime,gen ; dst_parent.mtime,gen ;
handle overwritten dst inode as above; if versioned bucket, update VIDX of both
names accordingly (a rename of a versioned object operates on the CURRENT
version; S3 has no native rename — this path is POSIX/rename only).
RENAME_EXCHANGE / RENAME_NOREPLACE (renameat2): EXCHANGE swaps two existing DENTs
atomically (both must exist); NOREPLACE fails with EEXIST if dst exists. Both
are single transactions.
22.4 unlink / rmdir and open-but-unlinked (POSIX orphan)
unlink(parent,name):
· target must not be a Directory (→ EISDIR; use rmdir).
· del DENT; target.nlink-- .
· if nlink == 0:
if open_count == 0 → delete INOD now + DELQ (OwnLayout) / CMTA.dead_bytes
(Slice).
if open_count > 0 → ORPHAN: keep INOD (no DENT), set a tombstone flag;
defer deletion to the final close (I17). Reads/writes
via the open handle continue to work.
open_count is client-session state tracked by the MDS via open/close RPCs (a
lightweight per-inode open table, NOT a durable lease). On MDS failover the
open table is rebuilt from client re-open/keepalive; an orphan with no live
opener after a grace period is reclaimed by the scrubber (handles crashed
clients).
rmdir(parent,name): target must be an EMPTY Directory (no DENT under its prefix,
checked by a 1-row range scan) → else ENOTEMPTY. del DENT + del INOD;
parent.nlink-- .
22.5 hardlink / symlink
hardlink(target_ino, parent, name): target MUST be a RegularFile (I18); never a
Directory. put DENT(parent,name)=target_ino; target.nlink++ ; target.ctime,gen.
Hard links share one inode (one body/Layout/Slice) → all names see the same data
and the same `gen`.
symlink(parent, name, target_path): put INOD(kind=Symlink, body=Inline(target));
put DENT. The target is an uninterpreted byte string resolved at lookup (§22.2);
dangling targets are allowed (POSIX).
22.6 readdir cursor (stable pagination)
readdir(dir, cursor, limit):
range-scan DENT keys in [0x02|dir|cursor.last_name+0x00 , 0x02|dir|0xFF…),
return up to `limit` (name, child_ino, kind_tag); next cursor = last name returned.
· The cursor is the last NAME (not an offset) → stable across concurrent inserts/
deletes (telldir/seekdir semantics): a create after the cursor will appear; a
name already passed will not reappear.
· readdirplus additionally returns the child InodeAttr (one extra get per entry,
or batched). Entries are name-ordered (lexicographic on UTF-8 bytes).
· S3 ListObjectsV2 maps marker/continuation-token to this name cursor;
delimiter='/' returns child Directory names as CommonPrefixes (a readdir that
stops at the delimiter level — §4).
22.7 atime / mtime / ctime rules
· mtime: data or directory-entry change (write, create/delete child).
· ctime: any metadata change (mode/owner/nlink/rename) OR data change.
· atime: read access; UPDATED LAZILY (relatime-style: only if atime < mtime or
older than a window) to avoid a metadata write per read — critical for the
read-heavy AI workload. Configurable (noatime default for dataset mounts).
23. S3 Object Semantics
23.1 Multipart upload — state machine & part↔chunk assembly
States: Initiated → Uploading → Completed | Aborted | Expired.
CreateMultipartUpload(B,key,meta):
allocate upload_id (u128); put MPUP{bucket_ino,key,started,init_meta,policy}.
No final inode yet — final offsets are unknown until Complete (they depend on
lower part_no sizes, which are finalized only at Complete, and parts may be
re-uploaded).
UploadPart(upload_id, part_no, data):
stage SELF-CONTAINED: allocate part_file_id (IALOC); write `data` as formula
chunks under part_file_id with MPUP.policy; put MPRT{size,etag=md5,checksums,
part_file_id,layout}. Re-upload of part_no replaces it (old part_file_id → DELQ).
CompleteMultipartUpload(upload_id, ordered [(part_no, etag)]):
validate etags, ordering, min part size (≥ 5 MiB except last).
final offset of part = Σ sizes of preceding parts (now known).
build body = Segmented([ Segment{offset,length,part_file_id,layout} … ])
— NO data copy, NO re-stripe (Complete is O(#parts) metadata).
allocate final inode; body = Segmented (or that part's Layout if exactly 1 part);
etag = "{md5(concat part md5s)}-{N}".
ONE txn: put final INOD; put DENT (+ implicit mkdir -p, §4); if versioned upsert
VIDX; delete MPUP + all MPRT. Staged part chunks become the object's chunks in
place (Segmented now owns the part_file_ids).
Part↔chunk alignment: because parts are staged self-contained and the object uses
a Segmented body, parts need NOT align to chunk_size — no re-stripe. A read at
object offset O binary-searches segments, then resolves within that segment's
part_file_id by formula (SPEC-PLACEMENT-001). Optional background re-stripe can
later collapse a hot Segmented object into one uniform Layout; not required.
23.2 UploadPartCopy / ListParts / Abort
UploadPartCopy(upload_id, part_no, src, range):
server-side: read src byte-range → stage as a part (as UploadPart) without
client round-trip. If src/dst share placement and the range is chunk-aligned,
the copy MAY be a metadata-only chunk reference (ref-counted CoW) — optimization;
default is a data copy.
ListParts(upload_id): range-scan MPRT over 0x0F|upload_id → (part_no,size,etag),
paginated by part_no cursor.
AbortMultipartUpload(upload_id): del MPUP + all MPRT; DELQ each staged
part_file_id → chunk GC. Idempotent.
23.3 Versioning & delete markers (refines §7 VIDX)
BUCK.versioning ∈ {disabled, enabled, suspended}.
PUT (enabled) : new inode + new version_id; prepend latest in VIDX;
DENT.child_ino → new inode (old versions retained).
PUT (disabled/susp.) : overwrite the single "null" version; old inode → unlink
(DELQ/CMTA if last ref).
DELETE current (en.) : insert DELETE MARKER VersionEntry (delete_marker=1,
inode=NIL) as latest; REMOVE the DENT (POSIX sees the name
gone; ListObjectsV2 omits it). GET current → DENT missing →
VIDX latest is a delete marker → 404 NoSuchKey. The key
still appears in ListObjectVersions.
DELETE version V : remove that VIDX entry; if its inode's refs reach 0 →
del INOD + DELQ/CMTA. Deleting the latest delete marker
reinstates the prior version → re-add DENT (becomes current).
GET version V : VIDX → entry.inode_id → INOD (bypasses DENT).
POSIX view : the FUSE mount sees the CURRENT non-delete-marker version;
version history is an S3-only concept via the S3 API.
23.4 Lifecycle (background worker over BUCK rules)
A lifecycle scanner (rustfs-scanner) walks each bucket and evaluates BUCK.lifecycle:
Expiration(days) : current older than days → DELETE (delete marker if
versioned; hard delete if not). Noncurrent-expiration →
drop old VIDX entries + GC their inodes/chunks.
Transition(days, tier) : move data to a remote tier; set SysMeta.transition;
free local chunks (DELQ) once the remote copy is durable;
GET of a transitioned object triggers restore (RestoreInfo).
AbortIncompleteMultipart(days) : MPUP older than days → Abort (§23.2).
Each action is a single metadata transaction (+ async data move for transition).
Timing is advisory / eventually-applied.
24. Sizing & Scale
InodeAttr 150–300 B · DentValue 20–40 B · ExtentEntry ~16 B · BucketConfig small.
1e9 files ≈ (250+30) B × 1e9 ≈ ~280 GB metadata + KV index → fits a distributed
KV. Billions of small files (AI datasets, node_modules, pip/conda) are normal:
their DATA is packed (CMTA), their METADATA is these rows. KV never holds bulk data.
25. Codec Versioning
Each value (except opaque XATR) = struct_version(u8=1) || MessagePack(struct).
Adding an OPTIONAL field with a valid None/zero default → no bump. Removing/
retyping a field or changing key layout → bump struct_version + ship a reader.
Greenfield: struct_version 1 is first and only at GA.
26. Conformance Checklist (M0/M1 exit)
□ All key encodings (§2) exact; LE/BE per table; readdir = range scan.
□ Namespace roots + reserved inodes pre-created; S3 key↔dir mapping with implicit
mkdir -p; ListObjects(delimiter) ↔ readdir; directory-marker keys.
□ InodeAttr + InodeKind + InodeBody(Empty/Inline/Slice/Layout/Segmented) round-trip.
□ Policy(EC|Replica) + Layout encode/decode; placement parity with chunk-store.
□ DentValue, VersionList, BucketConfig, EncryptionInfo, Volume, Snapshot,
ExtentEntry, ContainerMeta, ClusterMap, WriteIntent, DeleteTombstone, Multipart,
SegmentList, DirStat, Quota, Identity, AccessKey, LockEntry all round-trip.
□ Sparse / truncate / append / in-place semantics (§19); sparse reads zero.
□ Namespace semantics (§22): name rules; resolve incl. symlink ELOOP/ENOTDIR;
rename overwrite/dir-empty/loop-reject/EXCHANGE/NOREPLACE; unlink + open-orphan
(I17); rmdir-not-empty; hardlink-file-only; readdir name-cursor stability; relatime.
□ S3 semantics (§23): multipart Initiate/Upload/Complete(Segmented, no re-stripe)/
Abort/ListParts/UploadPartCopy; versioning + delete markers (DENT removal,
ListObjectVersions); lifecycle expire/transition/abort-incomplete.
□ Invariants I1–I22 enforced; nlink rules (files vs dirs).
□ Transaction groups (§21) are single SSI transactions; data-first ordering.
□ Id allocator: batched range reservation; file/container ids never collide.
□ struct_version gate rejects unknown versions.
Chinese mirror: rustfs-metadata-structures-ZH.md. Normative metadata companion to SPEC-CHUNK-001; consolidates MD-DESIGN-002/004/005 metadata definitions.