RustFS 文档 文档

RustFS Metadata Server — Change Plan

RustFS Metadata Server — Change Plan

Document:    RUSTFS-CHG-001
Status:      DRAFT
Version:     0.1.0
Companion:   RUSTFS-FEAT-001 (HLD), RUSTFS-LLD-001/002
Scope:       Migrating RustFS from embedded xl.meta metadata to a standalone,
             scalable Metadata Server (MDS) — incrementally, without breaking
             the running S3 object store.

0. The Core Problem

RustFS today has no metadata server. Metadata is stored as xl.meta files co-located with erasure-coded data shards on every disk, and accessed through ecstore + filemeta. There is no inode model, no directory tree, no centralized index.

The target (from RUSTFS-FEAT-001) is a standalone MDS with an inode model, RocksDB backend, Raft HA, leases, and a unified file/object/block namespace.

This is not a greenfield build — it is a migration of a live system. The change plan must:

  1. Never break the existing S3 API or stored data.
  2. Be incremental — each step ships independently and is reversible.
  3. Allow old (embedded) and new (MDS) metadata to coexist during transition.
  4. Provide a validated, online data migration path.

The strategy is the strangler fig pattern: insert a seam (a trait) between ecstore and its metadata operations, then grow the MDS behind that seam until it fully replaces the embedded path.


1. Current State Inventory (What We're Changing)

1.1 Where Metadata Lives Today

Component                         Role
────────────────────────────────  ──────────────────────────────────────────
filemeta crate (10K lines)        xl.meta on-disk format: FileMeta, FileInfo,
                                  ObjectPartInfo, ErasureInfo, versioning
                                  Header: [X,L,2,' '], XL_META_VERSION = 3

ecstore::store_api::ObjectInfo    In-memory object metadata representation

ecstore metadata_sys              Bucket + object metadata handling
(rustfs/src/storage/ecfs.rs)

init_bucket_metadata_sys          Startup step 8: loads bucket configs to memory
(rustfs/src/main.rs)

ecstore::set_disk (SetDisks)      Reads/writes xl.meta with quorum across the
                                  erasure set on every object op

ecstore::store_list_objects       ListObjects = prefix scan over xl.meta files

filemeta::metacache               In-memory metadata cache

lock crate (NamespaceLock)        Object-granularity locking

1.2 Current Metadata Operation Flow

S3 PutObject:
  app/object_usecase
    → storage/ecfs (validation, SSE)
      → ecstore SetDisks.put_object
        → write data shards (part.N) to N disks
        → write xl.meta to N disks (quorum)        ← METADATA WRITE

S3 GetObject:
  → ecstore SetDisks.get_object
    → read xl.meta from disks (quorum, pick latest) ← METADATA READ
    → read data shards, EC decode

S3 ListObjects:
  → ecstore walk: scan ALL xl.meta under prefix    ← O(N) METADATA SCAN

1.3 The Seams We Will Exploit

The good news: ecstore already abstracts metadata access behind internal
methods (read_xl_meta, write_xl_meta, list_path_entries). These are the
insertion points for the MetadataProvider trait. We do NOT need to rewrite
ecstore's data path — only its metadata calls.

2. Target State (Recap)

Standalone MDS process (rustfs-mds):
  ├── Inode model (InodeAttr, InodeId)         — RUSTFS-FEAT-001 §3.1
  ├── Directory tree (DirEntry)                 — §3.2
  ├── Chunk map (separate from inode)           — §3.3, DQ1 resolved below
  ├── RocksDB backend (CF: inodes/dentries/...) — §3.6
  ├── Raft HA (openraft)                        — §2.2
  ├── Lease + lock manager                      — §3.5
  └── gRPC services (Namespace/DataLayout/...)  — §4.1

ecstore becomes a DATA-ONLY engine:
  ├── chunk read/write (byte-range)             — RUSTFS-FEAT-001 §5
  ├── EC encode/decode (unchanged)
  └── no longer owns the namespace

2.1 Resolving DQ1: Chunk Map Storage

DQ1 (from HLD): store chunk map inline with inode, or in a separate CF?

DECISION: separate CF ("chunkmaps"), keyed by inode_id.

Rationale:
  - Small files (< 1 chunk): chunk map is tiny; the extra CF lookup is cheap
    and can be avoided by inlining the single-chunk case into the inode
    (a `Option<InlineChunk>` field on InodeAttr for files ≤ chunk_size).
  - Large files (millions of chunks): a 1 TB file = 262K chunk entries. Inlining
    that into the inode would bloat every stat() (which reads the inode).
    Separate CF keeps inode reads cheap; chunk map is fetched only on data I/O.
  - Hybrid: InodeAttr carries inline chunk locations for files ≤ 1 chunk;
    multi-chunk files use the chunkmaps CF.

3. Migration Strategy: Six Phases

Phase 0  Insert the seam (MetadataProvider trait)        — refactor, no behavior change
Phase 1  Build MDS as embedded library + standalone proc — new code, off by default
Phase 2  Dual-write + shadow-read validation             — correctness gate
Phase 3  Cutover for new data (MDS authoritative)        — behavior change, reversible
Phase 4  Migrate legacy xl.meta into MDS                  — online data migration
Phase 5  Inode model + unified namespace + HA             — full target state

Each phase is independently shippable and gated by acceptance criteria.


4. Phase 0 — Insert the Seam (Refactor Only)

Goal: introduce a MetadataProvider trait that wraps the current embedded behavior exactly, with zero functional change. This is the strangler-fig seam.

4.1 The Trait

// crate: rustfs-metaprovider (NEW, leaf crate)
// file:  crates/metaprovider/src/lib.rs

/// Abstraction over all namespace metadata operations.
///
/// Phase 0: the only implementation is `EmbeddedXlMetaProvider`, which
/// delegates to the existing ecstore xl.meta code paths. Behavior is identical.
///
/// Later phases add `MdsProvider` and `DualWriteProvider`.
#[async_trait]
pub trait MetadataProvider: Send + Sync {
    // --- Object/file metadata ---
    async fn get_object_meta(&self, bucket: &str, key: &str, version: Option<&str>)
        -> Result<ObjectMeta>;
    async fn put_object_meta(&self, meta: &ObjectMeta) -> Result<()>;
    async fn delete_object_meta(&self, bucket: &str, key: &str, version: Option<&str>)
        -> Result<()>;

    // --- Listing ---
    async fn list_objects(&self, bucket: &str, prefix: &str, opts: ListOpts)
        -> Result<ObjectListing>;

    // --- Bucket metadata ---
    async fn get_bucket_meta(&self, bucket: &str) -> Result<BucketMeta>;
    async fn put_bucket_meta(&self, meta: &BucketMeta) -> Result<()>;
    async fn list_buckets(&self) -> Result<Vec<BucketMeta>>;

    // --- Multipart ---
    async fn get_multipart(&self, upload_id: &str) -> Result<MultipartState>;
    async fn put_multipart(&self, state: &MultipartState) -> Result<()>;

    // --- Capabilities (lets callers know what's supported) ---
    fn capabilities(&self) -> MetaCapabilities;
}

/// Neutral metadata type — superset of what xl.meta and the inode model carry.
/// Maps 1:1 from ecstore::ObjectInfo today; gains inode fields later.
pub struct ObjectMeta {
    pub bucket:       String,
    pub key:          String,
    pub size:         u64,
    pub mod_time:     SystemTime,
    pub etag:         String,
    pub version_id:   Option<String>,
    pub parts:        Vec<PartMeta>,
    pub erasure:      ErasureMeta,
    pub user_meta:    BTreeMap<String, String>,
    // Phase 5: bridge to inode model
    pub inode:        Option<InodeId>,
}

4.2 The Embedded Implementation (Wraps Existing Code)

// crate: rustfs-metaprovider
// file:  crates/metaprovider/src/embedded.rs

/// Phase 0 implementation: delegates to the existing ecstore xl.meta path.
/// This is a thin adapter — it calls the same functions ecstore already uses,
/// just behind the trait. No logic moves; behavior is byte-for-byte identical.
pub struct EmbeddedXlMetaProvider {
    ecstore: Arc<ECStore>,
}

#[async_trait]
impl MetadataProvider for EmbeddedXlMetaProvider {
    async fn get_object_meta(&self, bucket: &str, key: &str, version: Option<&str>)
        -> Result<ObjectMeta>
    {
        // Call existing ecstore method, convert ObjectInfo → ObjectMeta
        let oi = self.ecstore.get_object_info(bucket, key, version).await?;
        Ok(ObjectMeta::from_object_info(oi))
    }

    async fn list_objects(&self, bucket: &str, prefix: &str, opts: ListOpts)
        -> Result<ObjectListing>
    {
        // Existing prefix-scan list (still O(N) — unchanged in Phase 0)
        let res = self.ecstore.list_objects(bucket, prefix, opts).await?;
        Ok(res.into())
    }
    // ... all other methods delegate to existing ecstore code ...

    fn capabilities(&self) -> MetaCapabilities {
        MetaCapabilities {
            inode_model: false,
            directory_index: false,   // listing is still O(N) scan
            byte_range_meta: false,
        }
    }
}

4.3 Wiring

// rustfs/src/storage/ecfs.rs — modify to call through the trait

// BEFORE:
//   let oi = self.ecstore.get_object_info(bucket, key, ver).await?;

// AFTER:
//   let meta = self.meta_provider.get_object_meta(bucket, key, ver).await?;

// The meta_provider is injected at startup. Phase 0 always uses
// EmbeddedXlMetaProvider, so behavior is unchanged.

4.4 Phase 0 Deliverables & Acceptance

Deliverables:
  ✓ crates/metaprovider/ with MetadataProvider trait + ObjectMeta types
  ✓ EmbeddedXlMetaProvider delegating to ecstore
  ✓ ecfs.rs and app/ layer call through the trait
  ✓ Conversion functions ObjectInfo ↔ ObjectMeta

Acceptance:
  ✓ Full existing S3 test suite passes unchanged (e2e_test crate)
  ✓ Zero performance regression (benchmark before/after)
  ✓ No new external dependencies

Risk: LOW (pure refactor). Rollback: revert the trait wiring.
Effort: ~3-4 weeks. This is the critical enabling step.

5. Phase 1 — Build the MDS (Off by Default)

Goal: implement rustfs-mds as both an embeddable library and a standalone process, implementing MetadataProvider. Not yet used in production paths.

5.1 Crate Structure

crates/mds/
├── src/
│   ├── lib.rs              # MdsProvider (implements MetadataProvider)
│   ├── store/
│   │   ├── mod.rs          # MetadataStore trait
│   │   ├── rocksdb.rs      # RocksDB backend (Phase 1: single node)
│   │   └── memory.rs       # in-memory backend (for tests)
│   ├── schema.rs           # CF definitions, key encoding
│   ├── inode_alloc.rs      # batched inode allocator
│   ├── service.rs          # gRPC service (standalone mode)
│   ├── bridge.rs           # ObjectMeta ↔ inode model conversion
│   └── main.rs             # rustfs-mds binary
└── Cargo.toml

5.2 Two Run Modes

Embedded mode (single-node RustFS):
  MDS runs as a library inside the rustfs process.
  RocksDB stored under {volume}/.rustfs.sys/mds/.
  No gRPC; direct in-process calls.

Standalone mode (clustered):
  rustfs-mds runs as a separate process / pod.
  rustfs processes connect via gRPC.
  Raft added in Phase 5.

5.3 MdsProvider Implements MetadataProvider

// crate: rustfs-mds
// file:  crates/mds/src/lib.rs

/// MDS-backed metadata provider. In Phase 1 it can store and retrieve metadata
/// but is not yet wired into production. Used in tests and shadow mode (Phase 2).
pub struct MdsProvider {
    store: Arc<dyn MetadataStore>,   // RocksDB
    cache: PathCache,
    alloc: InodeAllocator,
}

#[async_trait]
impl MetadataProvider for MdsProvider {
    async fn get_object_meta(&self, bucket: &str, key: &str, version: Option<&str>)
        -> Result<ObjectMeta>
    {
        // Resolve /buckets/{bucket}/{key} → inode (via path cache / dentry walk)
        let ino = self.resolve_s3(bucket, key).await?;
        let attr = self.store.get_inode(ino).await?;
        let chunks = self.store.get_chunk_map(ino).await?;
        Ok(self.bridge.to_object_meta(bucket, key, &attr, &chunks))
    }

    async fn list_objects(&self, bucket: &str, prefix: &str, opts: ListOpts)
        -> Result<ObjectListing>
    {
        // O(1)-per-directory readdir instead of O(N) scan!
        let dir = self.resolve_s3_prefix(bucket, prefix).await?;
        let listing = self.readdir(dir, opts.marker, opts.max_keys).await?;
        Ok(self.bridge.to_object_listing(listing, prefix, opts.delimiter))
    }

    fn capabilities(&self) -> MetaCapabilities {
        MetaCapabilities {
            inode_model: true,
            directory_index: true,    // O(1) listing
            byte_range_meta: true,
        }
    }
}

5.4 Phase 1 Deliverables & Acceptance

Deliverables:
  ✓ crates/mds/ with RocksDB store, schema, inode allocator
  ✓ MdsProvider implementing MetadataProvider
  ✓ rustfs-mds standalone binary + gRPC service
  ✓ Embedded mode integration (library)
  ✓ bridge.rs: ObjectMeta ↔ inode conversion

Acceptance:
  ✓ Unit + integration tests: create/get/list/delete via MdsProvider
  ✓ MdsProvider passes the SAME conformance test suite as EmbeddedXlMetaProvider
    (a shared test harness runs both against identical operation sequences)
  ✓ RocksDB crash recovery test (kill -9, restart, verify consistency)

Risk: LOW (new code, not in production path).
Effort: ~10-12 weeks.

6. Phase 2 — Dual-Write + Shadow-Read Validation

Goal: run both providers in parallel. Writes go to both; reads come from the authoritative embedded path but are also read from MDS and compared. This proves the MDS produces identical results before trusting it.

6.1 DualWriteProvider

// crate: rustfs-metaprovider
// file:  crates/metaprovider/src/dual.rs

/// Wraps two providers: `primary` (authoritative) and `shadow`.
///
/// Writes: applied to BOTH. If shadow write fails, log + metric but do NOT
///         fail the operation (primary is authoritative).
/// Reads:  served from primary. In validation mode, ALSO read from shadow
///         (async, off critical path) and compare; mismatches are logged
///         and counted as a correctness signal.
pub struct DualWriteProvider {
    primary: Arc<dyn MetadataProvider>,   // EmbeddedXlMetaProvider
    shadow:  Arc<dyn MetadataProvider>,   // MdsProvider
    mode:    DualMode,
    metrics: DualMetrics,
}

pub enum DualMode {
    /// Write both, read primary only (no comparison) — warm-up.
    WriteBoth,
    /// Write both, read both async + compare — validation.
    Validate,
}

#[async_trait]
impl MetadataProvider for DualWriteProvider {
    async fn put_object_meta(&self, meta: &ObjectMeta) -> Result<()> {
        // Primary first (authoritative)
        self.primary.put_object_meta(meta).await?;
        // Shadow best-effort
        if let Err(e) = self.shadow.put_object_meta(meta).await {
            self.metrics.shadow_write_errors.inc();
            warn!(error = %e, "shadow MDS write failed (non-fatal)");
        }
        Ok(())
    }

    async fn get_object_meta(&self, bucket: &str, key: &str, ver: Option<&str>)
        -> Result<ObjectMeta>
    {
        let primary = self.primary.get_object_meta(bucket, key, ver).await?;
        if matches!(self.mode, DualMode::Validate) {
            // Off critical path: compare shadow result
            let shadow = self.shadow.clone();
            let (b, k, expected) = (bucket.to_owned(), key.to_owned(), primary.clone());
            tokio::spawn(async move {
                match shadow.get_object_meta(&b, &k, ver).await {
                    Ok(s) if s.semantically_eq(&expected) => { /* match */ }
                    Ok(s) => metrics::record_mismatch(&b, &k, &expected, &s),
                    Err(e) => metrics::record_shadow_miss(&b, &k, e),
                }
            });
        }
        Ok(primary)
    }
}

6.2 Validation Metrics

Tracked continuously:
  mds_shadow_write_errors_total       — shadow writes that failed
  mds_shadow_read_mismatches_total    — reads where shadow ≠ primary
  mds_shadow_read_misses_total        — reads shadow couldn't find
  mds_shadow_list_mismatches_total    — listings that differed

Cutover gate (Phase 3) requires:
  - mismatch rate < 1 per 10^9 operations over a 2-week soak
  - zero unexplained mismatches (each must be root-caused)

6.3 Phase 2 Deliverables & Acceptance

Deliverables:
  ✓ DualWriteProvider with WriteBoth + Validate modes
  ✓ ObjectMeta::semantically_eq (ignores fields that legitimately differ,
    e.g., internal version representation)
  ✓ Mismatch reporting + dashboards
  ✓ Config flag: metadata.mode = "embedded" | "dual-write" | "dual-validate"

Acceptance:
  ✓ 2-week production soak in dual-validate on a canary cluster
  ✓ Mismatch rate below gate threshold
  ✓ All mismatches root-caused and fixed (typically: edge cases in versioning,
    delete markers, multipart assembly, special characters in keys)

Risk: LOW-MEDIUM (shadow path is non-authoritative; can't corrupt data).
  Watch: shadow write amplification (2x metadata writes) — monitor MDS load.
Effort: ~6-8 weeks including soak.

7. Phase 3 — Cutover (MDS Authoritative for New Data)

Goal: flip MDS to authoritative. Embedded xl.meta becomes the fallback / legacy reader. New writes are MDS-first.

7.1 Cutover Mechanics

Per-bucket cutover flag in MDS:
  bucket.metadata_backend = "embedded" | "mds"

New buckets: default to "mds".
Existing buckets: stay "embedded" until migrated (Phase 4).

Provider becomes a ROUTER:

  RoutingProvider.get_object_meta(bucket, key):
    match bucket_backend(bucket) {
      Mds      => mds.get_object_meta(...)          // authoritative
      Embedded => embedded.get_object_meta(...)      // legacy
      Migrating => {
        // during Phase 4 migration: MDS first, fall back to embedded
        match mds.get_object_meta(...) {
          Ok(m) => m,
          Err(NotFound) => embedded.get_object_meta(...),
        }
      }
    }
// crate: rustfs-metaprovider
// file:  crates/metaprovider/src/routing.rs

pub struct RoutingProvider {
    embedded: Arc<dyn MetadataProvider>,
    mds:      Arc<dyn MetadataProvider>,
    backends: Arc<BucketBackendMap>,   // bucket → backend, cached from MDS
}

#[async_trait]
impl MetadataProvider for RoutingProvider {
    async fn get_object_meta(&self, bucket: &str, key: &str, ver: Option<&str>)
        -> Result<ObjectMeta>
    {
        match self.backends.backend_for(bucket) {
            Backend::Mds => self.mds.get_object_meta(bucket, key, ver).await,
            Backend::Embedded => self.embedded.get_object_meta(bucket, key, ver).await,
            Backend::Migrating => {
                match self.mds.get_object_meta(bucket, key, ver).await {
                    Ok(m) => Ok(m),
                    Err(StorageError::PathNotFound(_)) =>
                        self.embedded.get_object_meta(bucket, key, ver).await,
                    Err(e) => Err(e),
                }
            }
        }
    }
}

7.2 Data Path Change: ecstore Stops Writing xl.meta for MDS Buckets

For MDS-backed buckets, ecstore writes ONLY data shards (part.N).
The xl.meta write is replaced by an MDS commit (chunk map + inode update).

ecstore SetDisks.put_object, for an MDS bucket:
  1. write data shards (part.N) — UNCHANGED
  2. SKIP xl.meta write
  3. return chunk locations to caller
  → caller (provider) commits chunk map + inode to MDS

This is the point where ecstore truly becomes data-only for new buckets.

7.3 Rollback Plan

If MDS misbehaves after cutover:
  - Flip bucket back to "embedded" — BUT only safe if xl.meta still written.
  - Therefore: in early Phase 3, keep dual-write ON (xl.meta still written as
    a safety net) even though MDS is authoritative for reads.
  - Once confidence is high (weeks of clean operation), disable the xl.meta
    safety net for MDS buckets to reclaim the write amplification.

Two sub-stages:
  3a. MDS authoritative for reads, xl.meta still written (reversible instantly)
  3b. MDS authoritative, xl.meta writes disabled (reversible only via migration)

7.4 Phase 3 Deliverables & Acceptance

Deliverables:
  ✓ RoutingProvider with per-bucket backend selection
  ✓ ecstore data-only mode for MDS buckets
  ✓ Admin API: set/get bucket metadata backend
  ✓ Stage 3a (safety net) → Stage 3b (net removed) toggle

Acceptance:
  ✓ New buckets serve all S3 ops via MDS, full conformance suite passes
  ✓ Listing latency drops dramatically (O(1) vs O(N)) — benchmark proof
  ✓ Rollback drill: flip a bucket MDS→embedded and back, verify no data loss
  ✓ 30-day soak on new-bucket workloads before Stage 3b

Risk: MEDIUM (MDS now authoritative). Mitigated by staged rollback + safety net.
Effort: ~8 weeks across both sub-stages.

8. Phase 4 — Migrate Legacy xl.meta Into MDS

Goal: convert existing embedded-backed buckets to MDS, online, without downtime.

8.1 Migration Engine

// crate: rustfs-mds
// file:  crates/mds/src/migrate.rs

/// Online migration of a bucket from embedded xl.meta to MDS.
///
/// Strategy: scan-and-build. Walk all xl.meta in the bucket, construct
/// inode + dentry + chunk map entries in MDS. The bucket is in "Migrating"
/// state throughout, so reads fall back to embedded for not-yet-migrated keys.
///
/// Writes during migration go to MDS (the bucket is already routed); the
/// scanner skips keys already present in MDS.
pub struct BucketMigrator {
    embedded: Arc<EmbeddedXlMetaProvider>,
    mds:      Arc<MdsProvider>,
    rate:     TokenBucket,        // throttle to protect production I/O
}

impl BucketMigrator {
    pub async fn migrate(&self, bucket: &str) -> Result<MigrationReport> {
        // 1. Set bucket state = Migrating (reads fall back to embedded)
        self.mds.set_bucket_backend(bucket, Backend::Migrating).await?;

        // 2. Ensure bucket directory inode exists in MDS
        let bucket_ino = self.mds.ensure_bucket_dir(bucket).await?;

        // 3. Scan all objects (reuse existing ecstore walk)
        let mut cursor = ListCursor::start();
        let mut report = MigrationReport::default();
        loop {
            let batch = self.embedded.list_objects(bucket, "", ListOpts {
                marker: cursor.marker(), max_keys: 1000, ..Default::default()
            }).await?;

            for obj in batch.objects {
                self.rate.acquire(1).await;
                // Skip if already in MDS (written during migration)
                if self.mds.exists(bucket, &obj.key).await? {
                    report.skipped += 1;
                    continue;
                }
                // Build inode + dentries (mkdir -p for path components)
                // + chunk map from the existing ErasureInfo/parts.
                // NO DATA IS MOVED — only metadata is translated.
                self.migrate_one(bucket_ino, bucket, &obj).await?;
                report.migrated += 1;
            }
            if !batch.has_more { break; }
            cursor = batch.next_cursor;
        }

        // 4. Flip to MDS-authoritative
        self.mds.set_bucket_backend(bucket, Backend::Mds).await?;
        Ok(report)
    }

    /// Translate one object's xl.meta into MDS inode + chunk map.
    /// Critically: the existing data shards (part.N) are NOT touched.
    /// The chunk map simply points at the existing shard locations.
    async fn migrate_one(&self, parent: InodeId, bucket: &str, obj: &ObjectMeta)
        -> Result<()>
    {
        // Create intermediate directories (mkdir -p)
        let parent_ino = self.mds.mkdir_p(parent, dirname(&obj.key)).await?;
        // Allocate inode
        let ino = self.mds.create_inode(parent_ino, basename(&obj.key),
                                        InodeKind::RegularFile, obj).await?;
        // Build chunk map from existing ErasureInfo — point at existing shards
        let chunk_map = self.build_chunk_map_from_erasure(ino, obj)?;
        self.mds.put_chunk_map(ino, &chunk_map).await?;
        Ok(())
    }
}

8.2 The Key Migration Insight

Migration translates METADATA ONLY. Data shards (part.N files) stay exactly
where they are. The MDS chunk map points at the existing shard locations.

  Before: xl.meta describes shards at /{disk}/{bucket}/{keyhash}/part.N
  After:  MDS chunk map describes the SAME shards at the SAME locations

This makes migration:
  - Fast (no data movement, just metadata translation)
  - Safe (data is never touched, so it can't be corrupted)
  - Reversible until the bucket is flipped to MDS-authoritative

Optional later step: re-stripe migrated data into the new chunk format
(for files that would benefit from the new striping). This is a separate,
lazy, opportunistic process — NOT part of the metadata migration.

8.3 Phase 4 Deliverables & Acceptance

Deliverables:
  ✓ BucketMigrator with online scan-and-build
  ✓ Chunk map construction from existing ErasureInfo
  ✓ Rate limiting to protect production
  ✓ Admin API: migrate bucket, migration status, pause/resume
  ✓ Verification tool: post-migration consistency check (MDS vs xl.meta)

Acceptance:
  ✓ Migrate a large bucket (100M+ objects) online with < 5% I/O impact
  ✓ Post-migration verification: 100% of objects readable via MDS path
  ✓ Crash-during-migration recovery (resume from cursor, idempotent)
  ✓ Concurrent writes during migration land correctly in MDS

Risk: MEDIUM (touches legacy data's metadata). Mitigated: data never moved,
  verification gate, reversible until flip.
Effort: ~8-10 weeks.

9. Phase 5 — Full Target State

Goal: complete the MDS to the RUSTFS-FEAT-001 design: inode model fully exposed (POSIX), HA via Raft, leases, sharding.

9.1 Sub-Phases

5a. Lease + lock manager (enables client caching coherency)
    → unblocks FUSE client and random-read optimizations (LLD-002)

5b. Raft HA (openraft) — MDS cluster, no single point of failure
    → all metadata mutations through Raft log
    → leases and locks become part of the replicated state machine

5c. MDS sharding (DNE — subtree partitioning)
    → namespace partitioned across MDS instances
    → cross-shard rename via 2PC

5d. Decommission embedded path
    → once all buckets are MDS-backed and stable, remove EmbeddedXlMetaProvider
      and the xl.meta write code from ecstore (keep xl.meta READER for any
      cold archival data, or migrate it all first)

9.2 Raft Integration Detail

// crate: rustfs-mds
// file:  crates/mds/src/raft.rs

/// All metadata MUTATIONS become Raft log entries. Reads can be:
///   - Linearizable (read from leader, after a no-op log commit), or
///   - Bounded-stale (read from any replica with a freshness bound)
///
/// The RocksDB store becomes the Raft state machine (apply log → mutate RocksDB).
pub enum MdsCommand {
    CreateInode(InodeAttr),
    SetInode(InodeId, InodeDelta),
    Unlink(InodeId, String),
    Rename { from: (InodeId, String), to: (InodeId, String) },
    PutChunkMap(InodeId, ChunkMap),
    GrantLease(Lease),
    RevokeLease(InodeId, ClientId),
    AcquireLock(ByteRangeLock),
    // ...
}

// Reads bypass Raft log (served from local RocksDB) when bounded-stale is OK;
// metadata-critical reads (e.g., during rename) go through the leader.

9.3 Phase 5 Deliverables & Acceptance

Deliverables:
  ✓ 5a: lease/lock manager + client revocation callbacks
  ✓ 5b: Raft HA MDS (3/5 node), failover < 2s
  ✓ 5c: subtree sharding + cross-shard 2PC rename
  ✓ 5d: embedded path removal (or archival-only reader)

Acceptance:
  ✓ POSIX FUSE mount fully functional (composes with FEAT-001 §6)
  ✓ MDS leader kill → automatic failover, no data loss, leases survive
  ✓ Sharded namespace scales metadata ops linearly with MDS count
  ✓ mdtest benchmark: target ops/s (see LLD-002 §9)

Risk: MEDIUM-HIGH (Raft correctness, sharding consistency).
  Mitigated by: openraft (proven library), extensive Jepsen-style testing.
Effort: ~6-9 months across sub-phases.

10. Cross-Cutting Concerns

10.1 Consistency During Transition

Concern: a write goes to MDS but the process crashes before data shards land
  (or vice versa) — metadata/data mismatch.

Resolution: commit ordering + reconciliation.
  Write order: data shards FIRST, then MDS commit.
    - If crash after data, before MDS commit: orphan chunks (GC reclaims them).
    - If crash after MDS commit: never happens (data already durable).
  A background reconciler (reuse rustfs-scanner) detects:
    - MDS entries with missing chunks → mark for heal or remove
    - Chunks with no MDS reference → orphans → GC after grace period

10.2 Performance Guardrails

Each phase has a perf gate vs the previous:
  Phase 0: zero regression (pure refactor)
  Phase 2: dual-write adds ≤ 2x metadata write cost (acceptable, temporary)
  Phase 3: MDS read path must be ≤ embedded latency; listing must be FASTER
  Phase 4: migration I/O impact ≤ 5% on production
  Phase 5: Raft commit latency ≤ target (WAL on fast device)

10.3 Observability for the Migration

Dashboards required from Phase 2 onward:
  - Provider routing breakdown (embedded vs MDS vs migrating ops/s)
  - Shadow mismatch rate (Phase 2 gate)
  - Per-bucket migration progress + ETA
  - MDS RocksDB metrics (compaction, write stalls, cache hit rate)
  - Metadata operation latency histograms by provider
  - Orphan chunk count (reconciler health)

10.4 Testing Strategy

Shared conformance harness:
  A single suite of metadata operation sequences runs against BOTH providers
  and asserts identical results. Run in CI from Phase 1 onward.

Chaos/fault injection (Phase 3+):
  - Kill MDS mid-write, mid-migration
  - Network partition between rustfs and MDS
  - RocksDB corruption injection → recovery
  - Raft (Phase 5): partition, leader kill, clock skew (Jepsen-style)

Migration testing (Phase 4):
  - Migrate while under concurrent read/write load
  - Resume after crash at every cursor position
  - Verify byte-for-byte equivalence post-migration

11. Sequencing & Dependencies

Phase  Name                         Effort     Depends on    Ships independently?
─────  ───────────────────────────  ─────────  ────────────  ────────────────────
0      Insert seam (trait)          3-4 wk     —             yes (no behavior change)
1      Build MDS (off)              10-12 wk   Phase 0       yes (dormant)
2      Dual-write + validate        6-8 wk     Phase 1       yes (canary)
3a     Cutover w/ safety net        4 wk       Phase 2       yes (per-bucket)
3b     Remove safety net            4 wk       Phase 3a soak yes (per-bucket)
4      Migrate legacy buckets       8-10 wk    Phase 3       yes (per-bucket)
5a     Leases/locks                 6-8 wk     Phase 3       yes
5b     Raft HA                      8-12 wk    Phase 5a      yes
5c     Sharding (DNE)               8-12 wk    Phase 5b      yes
5d     Decommission embedded        4 wk       Phase 4+5c    final step

Total: ~16-20 months to full target state, but VALUE SHIPS EARLY:
  - After Phase 3: new buckets get O(1) listing + inode model
  - After Phase 5a: POSIX FUSE + random-read optimizations work
  - Object storage NEVER breaks throughout

12. Decision Log

ID    Decision                                          Rationale
────  ────────────────────────────────────────────────  ──────────────────────
CD1   Strangler-fig via MetadataProvider trait          Incremental, reversible,
                                                         no big-bang rewrite

CD2   Chunk map in separate CF, inline single-chunk      Keeps stat() cheap for
                                                         huge files (DQ1 resolved)

CD3   Migration translates metadata only, never          Fast, safe, reversible;
      moves data shards                                  data can't be corrupted

CD4   Per-bucket cutover, not cluster-wide               Blast radius control,
                                                         gradual confidence

CD5   Dual-write safety net before removing xl.meta      Instant rollback during
                                                         the riskiest transition

CD6   data-shards-first, MDS-commit-second ordering      Crash → orphans (GC'able)
                                                         not dangling metadata

CD7   Reuse rustfs-scanner for reconciliation,           Don't rebuild proven
      rustfs-heal for repair                             subsystems

This change plan governs the metadata server migration. It is deliberately incremental: the existing S3 object store keeps working at every step, value ships after Phase 3, and full rollback is possible until Phase 3b. Update the decision log and per-phase acceptance status as implementation proceeds.