RustFS Metadata Server — Change Plan
RustFS Metadata Server — Change Plan
Document: RUSTFS-CHG-001
Status: DRAFT
Version: 0.1.0
Companion: RUSTFS-FEAT-001 (HLD), RUSTFS-LLD-001/002
Scope: Migrating RustFS from embedded xl.meta metadata to a standalone,
scalable Metadata Server (MDS) — incrementally, without breaking
the running S3 object store.
0. The Core Problem
RustFS today has no metadata server. Metadata is stored as xl.meta files
co-located with erasure-coded data shards on every disk, and accessed through
ecstore + filemeta. There is no inode model, no directory tree, no
centralized index.
The target (from RUSTFS-FEAT-001) is a standalone MDS with an inode model, RocksDB backend, Raft HA, leases, and a unified file/object/block namespace.
This is not a greenfield build — it is a migration of a live system. The change plan must:
- Never break the existing S3 API or stored data.
- Be incremental — each step ships independently and is reversible.
- Allow old (embedded) and new (MDS) metadata to coexist during transition.
- Provide a validated, online data migration path.
The strategy is the strangler fig pattern: insert a seam (a trait) between
ecstore and its metadata operations, then grow the MDS behind that seam until
it fully replaces the embedded path.
1. Current State Inventory (What We're Changing)
1.1 Where Metadata Lives Today
Component Role
──────────────────────────────── ──────────────────────────────────────────
filemeta crate (10K lines) xl.meta on-disk format: FileMeta, FileInfo,
ObjectPartInfo, ErasureInfo, versioning
Header: [X,L,2,' '], XL_META_VERSION = 3
ecstore::store_api::ObjectInfo In-memory object metadata representation
ecstore metadata_sys Bucket + object metadata handling
(rustfs/src/storage/ecfs.rs)
init_bucket_metadata_sys Startup step 8: loads bucket configs to memory
(rustfs/src/main.rs)
ecstore::set_disk (SetDisks) Reads/writes xl.meta with quorum across the
erasure set on every object op
ecstore::store_list_objects ListObjects = prefix scan over xl.meta files
filemeta::metacache In-memory metadata cache
lock crate (NamespaceLock) Object-granularity locking
1.2 Current Metadata Operation Flow
S3 PutObject:
app/object_usecase
→ storage/ecfs (validation, SSE)
→ ecstore SetDisks.put_object
→ write data shards (part.N) to N disks
→ write xl.meta to N disks (quorum) ← METADATA WRITE
S3 GetObject:
→ ecstore SetDisks.get_object
→ read xl.meta from disks (quorum, pick latest) ← METADATA READ
→ read data shards, EC decode
S3 ListObjects:
→ ecstore walk: scan ALL xl.meta under prefix ← O(N) METADATA SCAN
1.3 The Seams We Will Exploit
The good news: ecstore already abstracts metadata access behind internal
methods (read_xl_meta, write_xl_meta, list_path_entries). These are the
insertion points for the MetadataProvider trait. We do NOT need to rewrite
ecstore's data path — only its metadata calls.
2. Target State (Recap)
Standalone MDS process (rustfs-mds):
├── Inode model (InodeAttr, InodeId) — RUSTFS-FEAT-001 §3.1
├── Directory tree (DirEntry) — §3.2
├── Chunk map (separate from inode) — §3.3, DQ1 resolved below
├── RocksDB backend (CF: inodes/dentries/...) — §3.6
├── Raft HA (openraft) — §2.2
├── Lease + lock manager — §3.5
└── gRPC services (Namespace/DataLayout/...) — §4.1
ecstore becomes a DATA-ONLY engine:
├── chunk read/write (byte-range) — RUSTFS-FEAT-001 §5
├── EC encode/decode (unchanged)
└── no longer owns the namespace
2.1 Resolving DQ1: Chunk Map Storage
DQ1 (from HLD): store chunk map inline with inode, or in a separate CF?
DECISION: separate CF ("chunkmaps"), keyed by inode_id.
Rationale:
- Small files (< 1 chunk): chunk map is tiny; the extra CF lookup is cheap
and can be avoided by inlining the single-chunk case into the inode
(a `Option<InlineChunk>` field on InodeAttr for files ≤ chunk_size).
- Large files (millions of chunks): a 1 TB file = 262K chunk entries. Inlining
that into the inode would bloat every stat() (which reads the inode).
Separate CF keeps inode reads cheap; chunk map is fetched only on data I/O.
- Hybrid: InodeAttr carries inline chunk locations for files ≤ 1 chunk;
multi-chunk files use the chunkmaps CF.
3. Migration Strategy: Six Phases
Phase 0 Insert the seam (MetadataProvider trait) — refactor, no behavior change
Phase 1 Build MDS as embedded library + standalone proc — new code, off by default
Phase 2 Dual-write + shadow-read validation — correctness gate
Phase 3 Cutover for new data (MDS authoritative) — behavior change, reversible
Phase 4 Migrate legacy xl.meta into MDS — online data migration
Phase 5 Inode model + unified namespace + HA — full target state
Each phase is independently shippable and gated by acceptance criteria.
4. Phase 0 — Insert the Seam (Refactor Only)
Goal: introduce a MetadataProvider trait that wraps the current embedded
behavior exactly, with zero functional change. This is the strangler-fig seam.
4.1 The Trait
// crate: rustfs-metaprovider (NEW, leaf crate)
// file: crates/metaprovider/src/lib.rs
/// Abstraction over all namespace metadata operations.
///
/// Phase 0: the only implementation is `EmbeddedXlMetaProvider`, which
/// delegates to the existing ecstore xl.meta code paths. Behavior is identical.
///
/// Later phases add `MdsProvider` and `DualWriteProvider`.
#[async_trait]
pub trait MetadataProvider: Send + Sync {
// --- Object/file metadata ---
async fn get_object_meta(&self, bucket: &str, key: &str, version: Option<&str>)
-> Result<ObjectMeta>;
async fn put_object_meta(&self, meta: &ObjectMeta) -> Result<()>;
async fn delete_object_meta(&self, bucket: &str, key: &str, version: Option<&str>)
-> Result<()>;
// --- Listing ---
async fn list_objects(&self, bucket: &str, prefix: &str, opts: ListOpts)
-> Result<ObjectListing>;
// --- Bucket metadata ---
async fn get_bucket_meta(&self, bucket: &str) -> Result<BucketMeta>;
async fn put_bucket_meta(&self, meta: &BucketMeta) -> Result<()>;
async fn list_buckets(&self) -> Result<Vec<BucketMeta>>;
// --- Multipart ---
async fn get_multipart(&self, upload_id: &str) -> Result<MultipartState>;
async fn put_multipart(&self, state: &MultipartState) -> Result<()>;
// --- Capabilities (lets callers know what's supported) ---
fn capabilities(&self) -> MetaCapabilities;
}
/// Neutral metadata type — superset of what xl.meta and the inode model carry.
/// Maps 1:1 from ecstore::ObjectInfo today; gains inode fields later.
pub struct ObjectMeta {
pub bucket: String,
pub key: String,
pub size: u64,
pub mod_time: SystemTime,
pub etag: String,
pub version_id: Option<String>,
pub parts: Vec<PartMeta>,
pub erasure: ErasureMeta,
pub user_meta: BTreeMap<String, String>,
// Phase 5: bridge to inode model
pub inode: Option<InodeId>,
}
4.2 The Embedded Implementation (Wraps Existing Code)
// crate: rustfs-metaprovider
// file: crates/metaprovider/src/embedded.rs
/// Phase 0 implementation: delegates to the existing ecstore xl.meta path.
/// This is a thin adapter — it calls the same functions ecstore already uses,
/// just behind the trait. No logic moves; behavior is byte-for-byte identical.
pub struct EmbeddedXlMetaProvider {
ecstore: Arc<ECStore>,
}
#[async_trait]
impl MetadataProvider for EmbeddedXlMetaProvider {
async fn get_object_meta(&self, bucket: &str, key: &str, version: Option<&str>)
-> Result<ObjectMeta>
{
// Call existing ecstore method, convert ObjectInfo → ObjectMeta
let oi = self.ecstore.get_object_info(bucket, key, version).await?;
Ok(ObjectMeta::from_object_info(oi))
}
async fn list_objects(&self, bucket: &str, prefix: &str, opts: ListOpts)
-> Result<ObjectListing>
{
// Existing prefix-scan list (still O(N) — unchanged in Phase 0)
let res = self.ecstore.list_objects(bucket, prefix, opts).await?;
Ok(res.into())
}
// ... all other methods delegate to existing ecstore code ...
fn capabilities(&self) -> MetaCapabilities {
MetaCapabilities {
inode_model: false,
directory_index: false, // listing is still O(N) scan
byte_range_meta: false,
}
}
}
4.3 Wiring
// rustfs/src/storage/ecfs.rs — modify to call through the trait
// BEFORE:
// let oi = self.ecstore.get_object_info(bucket, key, ver).await?;
// AFTER:
// let meta = self.meta_provider.get_object_meta(bucket, key, ver).await?;
// The meta_provider is injected at startup. Phase 0 always uses
// EmbeddedXlMetaProvider, so behavior is unchanged.
4.4 Phase 0 Deliverables & Acceptance
Deliverables:
✓ crates/metaprovider/ with MetadataProvider trait + ObjectMeta types
✓ EmbeddedXlMetaProvider delegating to ecstore
✓ ecfs.rs and app/ layer call through the trait
✓ Conversion functions ObjectInfo ↔ ObjectMeta
Acceptance:
✓ Full existing S3 test suite passes unchanged (e2e_test crate)
✓ Zero performance regression (benchmark before/after)
✓ No new external dependencies
Risk: LOW (pure refactor). Rollback: revert the trait wiring.
Effort: ~3-4 weeks. This is the critical enabling step.
5. Phase 1 — Build the MDS (Off by Default)
Goal: implement rustfs-mds as both an embeddable library and a standalone
process, implementing MetadataProvider. Not yet used in production paths.
5.1 Crate Structure
crates/mds/
├── src/
│ ├── lib.rs # MdsProvider (implements MetadataProvider)
│ ├── store/
│ │ ├── mod.rs # MetadataStore trait
│ │ ├── rocksdb.rs # RocksDB backend (Phase 1: single node)
│ │ └── memory.rs # in-memory backend (for tests)
│ ├── schema.rs # CF definitions, key encoding
│ ├── inode_alloc.rs # batched inode allocator
│ ├── service.rs # gRPC service (standalone mode)
│ ├── bridge.rs # ObjectMeta ↔ inode model conversion
│ └── main.rs # rustfs-mds binary
└── Cargo.toml
5.2 Two Run Modes
Embedded mode (single-node RustFS):
MDS runs as a library inside the rustfs process.
RocksDB stored under {volume}/.rustfs.sys/mds/.
No gRPC; direct in-process calls.
Standalone mode (clustered):
rustfs-mds runs as a separate process / pod.
rustfs processes connect via gRPC.
Raft added in Phase 5.
5.3 MdsProvider Implements MetadataProvider
// crate: rustfs-mds
// file: crates/mds/src/lib.rs
/// MDS-backed metadata provider. In Phase 1 it can store and retrieve metadata
/// but is not yet wired into production. Used in tests and shadow mode (Phase 2).
pub struct MdsProvider {
store: Arc<dyn MetadataStore>, // RocksDB
cache: PathCache,
alloc: InodeAllocator,
}
#[async_trait]
impl MetadataProvider for MdsProvider {
async fn get_object_meta(&self, bucket: &str, key: &str, version: Option<&str>)
-> Result<ObjectMeta>
{
// Resolve /buckets/{bucket}/{key} → inode (via path cache / dentry walk)
let ino = self.resolve_s3(bucket, key).await?;
let attr = self.store.get_inode(ino).await?;
let chunks = self.store.get_chunk_map(ino).await?;
Ok(self.bridge.to_object_meta(bucket, key, &attr, &chunks))
}
async fn list_objects(&self, bucket: &str, prefix: &str, opts: ListOpts)
-> Result<ObjectListing>
{
// O(1)-per-directory readdir instead of O(N) scan!
let dir = self.resolve_s3_prefix(bucket, prefix).await?;
let listing = self.readdir(dir, opts.marker, opts.max_keys).await?;
Ok(self.bridge.to_object_listing(listing, prefix, opts.delimiter))
}
fn capabilities(&self) -> MetaCapabilities {
MetaCapabilities {
inode_model: true,
directory_index: true, // O(1) listing
byte_range_meta: true,
}
}
}
5.4 Phase 1 Deliverables & Acceptance
Deliverables:
✓ crates/mds/ with RocksDB store, schema, inode allocator
✓ MdsProvider implementing MetadataProvider
✓ rustfs-mds standalone binary + gRPC service
✓ Embedded mode integration (library)
✓ bridge.rs: ObjectMeta ↔ inode conversion
Acceptance:
✓ Unit + integration tests: create/get/list/delete via MdsProvider
✓ MdsProvider passes the SAME conformance test suite as EmbeddedXlMetaProvider
(a shared test harness runs both against identical operation sequences)
✓ RocksDB crash recovery test (kill -9, restart, verify consistency)
Risk: LOW (new code, not in production path).
Effort: ~10-12 weeks.
6. Phase 2 — Dual-Write + Shadow-Read Validation
Goal: run both providers in parallel. Writes go to both; reads come from the authoritative embedded path but are also read from MDS and compared. This proves the MDS produces identical results before trusting it.
6.1 DualWriteProvider
// crate: rustfs-metaprovider
// file: crates/metaprovider/src/dual.rs
/// Wraps two providers: `primary` (authoritative) and `shadow`.
///
/// Writes: applied to BOTH. If shadow write fails, log + metric but do NOT
/// fail the operation (primary is authoritative).
/// Reads: served from primary. In validation mode, ALSO read from shadow
/// (async, off critical path) and compare; mismatches are logged
/// and counted as a correctness signal.
pub struct DualWriteProvider {
primary: Arc<dyn MetadataProvider>, // EmbeddedXlMetaProvider
shadow: Arc<dyn MetadataProvider>, // MdsProvider
mode: DualMode,
metrics: DualMetrics,
}
pub enum DualMode {
/// Write both, read primary only (no comparison) — warm-up.
WriteBoth,
/// Write both, read both async + compare — validation.
Validate,
}
#[async_trait]
impl MetadataProvider for DualWriteProvider {
async fn put_object_meta(&self, meta: &ObjectMeta) -> Result<()> {
// Primary first (authoritative)
self.primary.put_object_meta(meta).await?;
// Shadow best-effort
if let Err(e) = self.shadow.put_object_meta(meta).await {
self.metrics.shadow_write_errors.inc();
warn!(error = %e, "shadow MDS write failed (non-fatal)");
}
Ok(())
}
async fn get_object_meta(&self, bucket: &str, key: &str, ver: Option<&str>)
-> Result<ObjectMeta>
{
let primary = self.primary.get_object_meta(bucket, key, ver).await?;
if matches!(self.mode, DualMode::Validate) {
// Off critical path: compare shadow result
let shadow = self.shadow.clone();
let (b, k, expected) = (bucket.to_owned(), key.to_owned(), primary.clone());
tokio::spawn(async move {
match shadow.get_object_meta(&b, &k, ver).await {
Ok(s) if s.semantically_eq(&expected) => { /* match */ }
Ok(s) => metrics::record_mismatch(&b, &k, &expected, &s),
Err(e) => metrics::record_shadow_miss(&b, &k, e),
}
});
}
Ok(primary)
}
}
6.2 Validation Metrics
Tracked continuously:
mds_shadow_write_errors_total — shadow writes that failed
mds_shadow_read_mismatches_total — reads where shadow ≠ primary
mds_shadow_read_misses_total — reads shadow couldn't find
mds_shadow_list_mismatches_total — listings that differed
Cutover gate (Phase 3) requires:
- mismatch rate < 1 per 10^9 operations over a 2-week soak
- zero unexplained mismatches (each must be root-caused)
6.3 Phase 2 Deliverables & Acceptance
Deliverables:
✓ DualWriteProvider with WriteBoth + Validate modes
✓ ObjectMeta::semantically_eq (ignores fields that legitimately differ,
e.g., internal version representation)
✓ Mismatch reporting + dashboards
✓ Config flag: metadata.mode = "embedded" | "dual-write" | "dual-validate"
Acceptance:
✓ 2-week production soak in dual-validate on a canary cluster
✓ Mismatch rate below gate threshold
✓ All mismatches root-caused and fixed (typically: edge cases in versioning,
delete markers, multipart assembly, special characters in keys)
Risk: LOW-MEDIUM (shadow path is non-authoritative; can't corrupt data).
Watch: shadow write amplification (2x metadata writes) — monitor MDS load.
Effort: ~6-8 weeks including soak.
7. Phase 3 — Cutover (MDS Authoritative for New Data)
Goal: flip MDS to authoritative. Embedded xl.meta becomes the fallback / legacy reader. New writes are MDS-first.
7.1 Cutover Mechanics
Per-bucket cutover flag in MDS:
bucket.metadata_backend = "embedded" | "mds"
New buckets: default to "mds".
Existing buckets: stay "embedded" until migrated (Phase 4).
Provider becomes a ROUTER:
RoutingProvider.get_object_meta(bucket, key):
match bucket_backend(bucket) {
Mds => mds.get_object_meta(...) // authoritative
Embedded => embedded.get_object_meta(...) // legacy
Migrating => {
// during Phase 4 migration: MDS first, fall back to embedded
match mds.get_object_meta(...) {
Ok(m) => m,
Err(NotFound) => embedded.get_object_meta(...),
}
}
}
// crate: rustfs-metaprovider
// file: crates/metaprovider/src/routing.rs
pub struct RoutingProvider {
embedded: Arc<dyn MetadataProvider>,
mds: Arc<dyn MetadataProvider>,
backends: Arc<BucketBackendMap>, // bucket → backend, cached from MDS
}
#[async_trait]
impl MetadataProvider for RoutingProvider {
async fn get_object_meta(&self, bucket: &str, key: &str, ver: Option<&str>)
-> Result<ObjectMeta>
{
match self.backends.backend_for(bucket) {
Backend::Mds => self.mds.get_object_meta(bucket, key, ver).await,
Backend::Embedded => self.embedded.get_object_meta(bucket, key, ver).await,
Backend::Migrating => {
match self.mds.get_object_meta(bucket, key, ver).await {
Ok(m) => Ok(m),
Err(StorageError::PathNotFound(_)) =>
self.embedded.get_object_meta(bucket, key, ver).await,
Err(e) => Err(e),
}
}
}
}
}
7.2 Data Path Change: ecstore Stops Writing xl.meta for MDS Buckets
For MDS-backed buckets, ecstore writes ONLY data shards (part.N).
The xl.meta write is replaced by an MDS commit (chunk map + inode update).
ecstore SetDisks.put_object, for an MDS bucket:
1. write data shards (part.N) — UNCHANGED
2. SKIP xl.meta write
3. return chunk locations to caller
→ caller (provider) commits chunk map + inode to MDS
This is the point where ecstore truly becomes data-only for new buckets.
7.3 Rollback Plan
If MDS misbehaves after cutover:
- Flip bucket back to "embedded" — BUT only safe if xl.meta still written.
- Therefore: in early Phase 3, keep dual-write ON (xl.meta still written as
a safety net) even though MDS is authoritative for reads.
- Once confidence is high (weeks of clean operation), disable the xl.meta
safety net for MDS buckets to reclaim the write amplification.
Two sub-stages:
3a. MDS authoritative for reads, xl.meta still written (reversible instantly)
3b. MDS authoritative, xl.meta writes disabled (reversible only via migration)
7.4 Phase 3 Deliverables & Acceptance
Deliverables:
✓ RoutingProvider with per-bucket backend selection
✓ ecstore data-only mode for MDS buckets
✓ Admin API: set/get bucket metadata backend
✓ Stage 3a (safety net) → Stage 3b (net removed) toggle
Acceptance:
✓ New buckets serve all S3 ops via MDS, full conformance suite passes
✓ Listing latency drops dramatically (O(1) vs O(N)) — benchmark proof
✓ Rollback drill: flip a bucket MDS→embedded and back, verify no data loss
✓ 30-day soak on new-bucket workloads before Stage 3b
Risk: MEDIUM (MDS now authoritative). Mitigated by staged rollback + safety net.
Effort: ~8 weeks across both sub-stages.
8. Phase 4 — Migrate Legacy xl.meta Into MDS
Goal: convert existing embedded-backed buckets to MDS, online, without downtime.
8.1 Migration Engine
// crate: rustfs-mds
// file: crates/mds/src/migrate.rs
/// Online migration of a bucket from embedded xl.meta to MDS.
///
/// Strategy: scan-and-build. Walk all xl.meta in the bucket, construct
/// inode + dentry + chunk map entries in MDS. The bucket is in "Migrating"
/// state throughout, so reads fall back to embedded for not-yet-migrated keys.
///
/// Writes during migration go to MDS (the bucket is already routed); the
/// scanner skips keys already present in MDS.
pub struct BucketMigrator {
embedded: Arc<EmbeddedXlMetaProvider>,
mds: Arc<MdsProvider>,
rate: TokenBucket, // throttle to protect production I/O
}
impl BucketMigrator {
pub async fn migrate(&self, bucket: &str) -> Result<MigrationReport> {
// 1. Set bucket state = Migrating (reads fall back to embedded)
self.mds.set_bucket_backend(bucket, Backend::Migrating).await?;
// 2. Ensure bucket directory inode exists in MDS
let bucket_ino = self.mds.ensure_bucket_dir(bucket).await?;
// 3. Scan all objects (reuse existing ecstore walk)
let mut cursor = ListCursor::start();
let mut report = MigrationReport::default();
loop {
let batch = self.embedded.list_objects(bucket, "", ListOpts {
marker: cursor.marker(), max_keys: 1000, ..Default::default()
}).await?;
for obj in batch.objects {
self.rate.acquire(1).await;
// Skip if already in MDS (written during migration)
if self.mds.exists(bucket, &obj.key).await? {
report.skipped += 1;
continue;
}
// Build inode + dentries (mkdir -p for path components)
// + chunk map from the existing ErasureInfo/parts.
// NO DATA IS MOVED — only metadata is translated.
self.migrate_one(bucket_ino, bucket, &obj).await?;
report.migrated += 1;
}
if !batch.has_more { break; }
cursor = batch.next_cursor;
}
// 4. Flip to MDS-authoritative
self.mds.set_bucket_backend(bucket, Backend::Mds).await?;
Ok(report)
}
/// Translate one object's xl.meta into MDS inode + chunk map.
/// Critically: the existing data shards (part.N) are NOT touched.
/// The chunk map simply points at the existing shard locations.
async fn migrate_one(&self, parent: InodeId, bucket: &str, obj: &ObjectMeta)
-> Result<()>
{
// Create intermediate directories (mkdir -p)
let parent_ino = self.mds.mkdir_p(parent, dirname(&obj.key)).await?;
// Allocate inode
let ino = self.mds.create_inode(parent_ino, basename(&obj.key),
InodeKind::RegularFile, obj).await?;
// Build chunk map from existing ErasureInfo — point at existing shards
let chunk_map = self.build_chunk_map_from_erasure(ino, obj)?;
self.mds.put_chunk_map(ino, &chunk_map).await?;
Ok(())
}
}
8.2 The Key Migration Insight
Migration translates METADATA ONLY. Data shards (part.N files) stay exactly
where they are. The MDS chunk map points at the existing shard locations.
Before: xl.meta describes shards at /{disk}/{bucket}/{keyhash}/part.N
After: MDS chunk map describes the SAME shards at the SAME locations
This makes migration:
- Fast (no data movement, just metadata translation)
- Safe (data is never touched, so it can't be corrupted)
- Reversible until the bucket is flipped to MDS-authoritative
Optional later step: re-stripe migrated data into the new chunk format
(for files that would benefit from the new striping). This is a separate,
lazy, opportunistic process — NOT part of the metadata migration.
8.3 Phase 4 Deliverables & Acceptance
Deliverables:
✓ BucketMigrator with online scan-and-build
✓ Chunk map construction from existing ErasureInfo
✓ Rate limiting to protect production
✓ Admin API: migrate bucket, migration status, pause/resume
✓ Verification tool: post-migration consistency check (MDS vs xl.meta)
Acceptance:
✓ Migrate a large bucket (100M+ objects) online with < 5% I/O impact
✓ Post-migration verification: 100% of objects readable via MDS path
✓ Crash-during-migration recovery (resume from cursor, idempotent)
✓ Concurrent writes during migration land correctly in MDS
Risk: MEDIUM (touches legacy data's metadata). Mitigated: data never moved,
verification gate, reversible until flip.
Effort: ~8-10 weeks.
9. Phase 5 — Full Target State
Goal: complete the MDS to the RUSTFS-FEAT-001 design: inode model fully exposed (POSIX), HA via Raft, leases, sharding.
9.1 Sub-Phases
5a. Lease + lock manager (enables client caching coherency)
→ unblocks FUSE client and random-read optimizations (LLD-002)
5b. Raft HA (openraft) — MDS cluster, no single point of failure
→ all metadata mutations through Raft log
→ leases and locks become part of the replicated state machine
5c. MDS sharding (DNE — subtree partitioning)
→ namespace partitioned across MDS instances
→ cross-shard rename via 2PC
5d. Decommission embedded path
→ once all buckets are MDS-backed and stable, remove EmbeddedXlMetaProvider
and the xl.meta write code from ecstore (keep xl.meta READER for any
cold archival data, or migrate it all first)
9.2 Raft Integration Detail
// crate: rustfs-mds
// file: crates/mds/src/raft.rs
/// All metadata MUTATIONS become Raft log entries. Reads can be:
/// - Linearizable (read from leader, after a no-op log commit), or
/// - Bounded-stale (read from any replica with a freshness bound)
///
/// The RocksDB store becomes the Raft state machine (apply log → mutate RocksDB).
pub enum MdsCommand {
CreateInode(InodeAttr),
SetInode(InodeId, InodeDelta),
Unlink(InodeId, String),
Rename { from: (InodeId, String), to: (InodeId, String) },
PutChunkMap(InodeId, ChunkMap),
GrantLease(Lease),
RevokeLease(InodeId, ClientId),
AcquireLock(ByteRangeLock),
// ...
}
// Reads bypass Raft log (served from local RocksDB) when bounded-stale is OK;
// metadata-critical reads (e.g., during rename) go through the leader.
9.3 Phase 5 Deliverables & Acceptance
Deliverables:
✓ 5a: lease/lock manager + client revocation callbacks
✓ 5b: Raft HA MDS (3/5 node), failover < 2s
✓ 5c: subtree sharding + cross-shard 2PC rename
✓ 5d: embedded path removal (or archival-only reader)
Acceptance:
✓ POSIX FUSE mount fully functional (composes with FEAT-001 §6)
✓ MDS leader kill → automatic failover, no data loss, leases survive
✓ Sharded namespace scales metadata ops linearly with MDS count
✓ mdtest benchmark: target ops/s (see LLD-002 §9)
Risk: MEDIUM-HIGH (Raft correctness, sharding consistency).
Mitigated by: openraft (proven library), extensive Jepsen-style testing.
Effort: ~6-9 months across sub-phases.
10. Cross-Cutting Concerns
10.1 Consistency During Transition
Concern: a write goes to MDS but the process crashes before data shards land
(or vice versa) — metadata/data mismatch.
Resolution: commit ordering + reconciliation.
Write order: data shards FIRST, then MDS commit.
- If crash after data, before MDS commit: orphan chunks (GC reclaims them).
- If crash after MDS commit: never happens (data already durable).
A background reconciler (reuse rustfs-scanner) detects:
- MDS entries with missing chunks → mark for heal or remove
- Chunks with no MDS reference → orphans → GC after grace period
10.2 Performance Guardrails
Each phase has a perf gate vs the previous:
Phase 0: zero regression (pure refactor)
Phase 2: dual-write adds ≤ 2x metadata write cost (acceptable, temporary)
Phase 3: MDS read path must be ≤ embedded latency; listing must be FASTER
Phase 4: migration I/O impact ≤ 5% on production
Phase 5: Raft commit latency ≤ target (WAL on fast device)
10.3 Observability for the Migration
Dashboards required from Phase 2 onward:
- Provider routing breakdown (embedded vs MDS vs migrating ops/s)
- Shadow mismatch rate (Phase 2 gate)
- Per-bucket migration progress + ETA
- MDS RocksDB metrics (compaction, write stalls, cache hit rate)
- Metadata operation latency histograms by provider
- Orphan chunk count (reconciler health)
10.4 Testing Strategy
Shared conformance harness:
A single suite of metadata operation sequences runs against BOTH providers
and asserts identical results. Run in CI from Phase 1 onward.
Chaos/fault injection (Phase 3+):
- Kill MDS mid-write, mid-migration
- Network partition between rustfs and MDS
- RocksDB corruption injection → recovery
- Raft (Phase 5): partition, leader kill, clock skew (Jepsen-style)
Migration testing (Phase 4):
- Migrate while under concurrent read/write load
- Resume after crash at every cursor position
- Verify byte-for-byte equivalence post-migration
11. Sequencing & Dependencies
Phase Name Effort Depends on Ships independently?
───── ─────────────────────────── ───────── ──────────── ────────────────────
0 Insert seam (trait) 3-4 wk — yes (no behavior change)
1 Build MDS (off) 10-12 wk Phase 0 yes (dormant)
2 Dual-write + validate 6-8 wk Phase 1 yes (canary)
3a Cutover w/ safety net 4 wk Phase 2 yes (per-bucket)
3b Remove safety net 4 wk Phase 3a soak yes (per-bucket)
4 Migrate legacy buckets 8-10 wk Phase 3 yes (per-bucket)
5a Leases/locks 6-8 wk Phase 3 yes
5b Raft HA 8-12 wk Phase 5a yes
5c Sharding (DNE) 8-12 wk Phase 5b yes
5d Decommission embedded 4 wk Phase 4+5c final step
Total: ~16-20 months to full target state, but VALUE SHIPS EARLY:
- After Phase 3: new buckets get O(1) listing + inode model
- After Phase 5a: POSIX FUSE + random-read optimizations work
- Object storage NEVER breaks throughout
12. Decision Log
ID Decision Rationale
──── ──────────────────────────────────────────────── ──────────────────────
CD1 Strangler-fig via MetadataProvider trait Incremental, reversible,
no big-bang rewrite
CD2 Chunk map in separate CF, inline single-chunk Keeps stat() cheap for
huge files (DQ1 resolved)
CD3 Migration translates metadata only, never Fast, safe, reversible;
moves data shards data can't be corrupted
CD4 Per-bucket cutover, not cluster-wide Blast radius control,
gradual confidence
CD5 Dual-write safety net before removing xl.meta Instant rollback during
the riskiest transition
CD6 data-shards-first, MDS-commit-second ordering Crash → orphans (GC'able)
not dangling metadata
CD7 Reuse rustfs-scanner for reconciliation, Don't rebuild proven
rustfs-heal for repair subsystems
This change plan governs the metadata server migration. It is deliberately incremental: the existing S3 object store keeps working at every step, value ships after Phase 3, and full rollback is possible until Phase 3b. Update the decision log and per-phase acceptance status as implementation proceeds.