RustFS 文档 文档

RustFS Unified Storage Platform

RustFS Unified Storage Platform

Feature Specification & High-Level Design

Document:    RUSTFS-FEAT-001
Status:      DRAFT
Version:     0.1.0
Authors:     Architecture Team
Created:     2026-06-05
Audience:    Engineering (implementation-ready)
Codebase:    github.com/rustfs/rustfs

0. Document Purpose

This specification defines the design of RustFS Unified Storage — a single storage platform that exposes file (POSIX), object (S3), and block access to the same data through one namespace. It is written for engineers who will implement the system. Every section maps to concrete crates, types, protobuf definitions, and code paths in the RustFS codebase.

The target workloads are:

  • AI training / inference: POSIX open → fseek → fread from hundreds of GPU nodes reading the same dataset concurrently, with high aggregate bandwidth (100+ GB/s cluster) and low per-op latency (<1 ms metadata, <5 ms first-byte data).
  • VM / bare-metal boot disks: 4K random read/write block I/O over iSCSI or NVMe-oF, with sub-millisecond latency, thin provisioning, and COW snapshots.
  • Object storage: existing S3-compatible API, unchanged, fully interoperable with the file and block namespaces.

1. Unified Namespace Model

1.1 Design Principle: One Namespace, Three Projections

Every piece of data in the cluster lives under a single hierarchical namespace rooted at /. The three access protocols project different views of this namespace:

                        ┌──────────────────────┐
                        │   Unified Namespace   │
                        │         /             │
                        │   ┌─────┴──────┐      │
                        │   │  volumes/  │ ...  │
                        │   │  data/     │      │
                        │   │  buckets/  │      │
                        │   └────────────┘      │
                        └──────┬───┬───┬────────┘
                               │   │   │
               ┌───────────────┘   │   └───────────────┐
               ▼                   ▼                   ▼
        ┌─────────────┐   ┌──────────────┐   ┌──────────────┐
        │  POSIX View  │   │   S3 View    │   │  Block View  │
        │              │   │              │   │              │
        │ mount -t     │   │ GET /bucket  │   │ /dev/rustfs/ │
        │  rustfs      │   │ /key HTTP/1.1│   │  vol-001     │
        │  /mnt/rustfs │   │              │   │              │
        │              │   │ bucket = dir  │   │ volume =     │
        │ ls, cat,     │   │ key = path   │   │  extent-     │
        │ fseek, fread │   │ under bucket │   │  mapped file │
        └─────────────┘   └──────────────┘   └──────────────┘

1.2 Namespace Hierarchy

/                                    # root inode (ino=1)
├── buckets/                         # S3 bucket mount root
│   ├── training-data/               # S3 bucket → directory
│   │   ├── imagenet/
│   │   │   ├── shard-0001.tar       # regular file (S3 object)
│   │   │   └── shard-0002.tar
│   │   └── llm/
│   │       └── tokens.bin
│   └── model-checkpoints/
│       └── epoch-042/
│           └── weights.safetensors
├── volumes/                         # block volume mount root
│   ├── vm-disk-001                  # block device special file
│   ├── vm-disk-002
│   └── db-wal-001
└── scratch/                         # general POSIX workspace
    └── user-alice/
        └── experiment-7/

1.3 Identity Mapping Rules

Every entity in the namespace is an inode. The access protocol determines how the inode is presented:

Inode kind          POSIX view              S3 view                  Block view
─────────────────── ─────────────────────── ──────────────────────── ──────────────
RegularFile         regular file            object (bucket = parent  N/A
                    (open/read/write/seek)  dir under /buckets/)

Directory           directory               bucket (if depth=1       N/A
                    (readdir/mkdir/rmdir)   under /buckets/) or
                                            common prefix

Symlink             symlink                 not exposed              N/A

BlockDevice         /dev/rustfs/vol-xxx     not exposed              LUN / NVMe
                    (can be opened with                              namespace
                    O_RDWR for raw I/O)

1.4 Path ↔ S3 Key Translation

POSIX path:  /buckets/training-data/imagenet/shard-0001.tar
S3 address:  bucket = "training-data", key = "imagenet/shard-0001.tar"

Rule:
  bucket_name = first path component under /buckets/
  object_key  = remaining path components joined by "/"

Reverse (S3 → POSIX):
  S3 PUT training-data/new-folder/file.bin
    → MDS: mkdir_p(/buckets/training-data/new-folder/)
    → MDS: create(/buckets/training-data/new-folder/file.bin)

2. System Architecture

2.1 Component Topology

┌─────────────────────────────────────────────────────────────────────┐
│                         CLIENT TIER                                  │
│                                                                      │
│  ┌──────────┐  ┌──────────┐  ┌──────────┐  ┌──────────────────┐    │
│  │  FUSE    │  │  S3 SDK  │  │  iSCSI   │  │  NVMe-oF         │    │
│  │  Client  │  │  Client  │  │  Initiator│  │  Host Driver     │    │
│  │(rustfs-  │  │(any S3   │  │(Linux    │  │(kernel nvme-tcp  │    │
│  │ mount)   │  │ client)  │  │ open-    │  │ or nvme-rdma)    │    │
│  │          │  │          │  │ iscsi)   │  │                  │    │
│  └────┬─────┘  └────┬─────┘  └────┬─────┘  └────────┬─────────┘    │
│       │              │              │                 │              │
│       │ gRPC         │ HTTP         │ iSCSI           │ NVMe-oF     │
└───────┼──────────────┼──────────────┼─────────────────┼─────────────┘
        │              │              │                 │
═══════════════════════════════════════════════════════════════════════
        │              │              │                 │
┌───────┼──────────────┼──────────────┼─────────────────┼─────────────┐
│       │         GATEWAY TIER        │                 │              │
│       │              │              │                 │              │
│       │         ┌────┴─────┐   ┌────┴────┐    ┌──────┴──────┐      │
│       │         │  S3 API  │   │  iSCSI  │    │  NVMe-oF    │      │
│       │         │  Server  │   │  Target │    │  Target     │      │
│       │         │ (existing│   │(rustfs- │    │(rustfs-     │      │
│       │         │  rustfs  │   │ iscsi)  │    │ nvmeof)     │      │
│       │         │  binary) │   │         │    │             │      │
│       │         └────┬─────┘   └────┬────┘    └──────┬──────┘      │
│       │              │              │                 │              │
└───────┼──────────────┼──────────────┼─────────────────┼─────────────┘
        │              │              │                 │
        ▼              ▼              ▼                 ▼
┌─────────────────────────────────────────────────────────────────────┐
│                      METADATA TIER (MDS)                             │
│                                                                      │
│  ┌────────────────────────────────────────────────────────────────┐  │
│  │                    MDS Cluster (Raft)                          │  │
│  │                                                                │  │
│  │  ┌──────────┐  ┌──────────┐  ┌──────────┐                    │  │
│  │  │  MDS-0   │──│  MDS-1   │──│  MDS-2   │   (3 or 5 nodes)  │  │
│  │  │ (leader) │  │(follower)│  │(follower)│                    │  │
│  │  └──────────┘  └──────────┘  └──────────┘                    │  │
│  │       │                                                       │  │
│  │  ┌────┴────────────────────────────────────────────┐          │  │
│  │  │  RocksDB: inodes, dentries, chunk maps,         │          │  │
│  │  │           volumes, leases, locks                 │          │  │
│  │  └─────────────────────────────────────────────────┘          │  │
│  └────────────────────────────────────────────────────────────────┘  │
│                                                                      │
└─────────────────────────────────────────────────────────────────────┘
        │
        ▼
┌─────────────────────────────────────────────────────────────────────┐
│                      DATA TIER (DSS — Data Storage Servers)          │
│                                                                      │
│  ┌─────────┐  ┌─────────┐  ┌─────────┐         ┌─────────┐        │
│  │  DSS-0  │  │  DSS-1  │  │  DSS-2  │  ...    │  DSS-N  │        │
│  │         │  │         │  │         │         │         │        │
│  │ chunk   │  │ chunk   │  │ chunk   │         │ chunk   │        │
│  │ service │  │ service │  │ service │         │ service │        │
│  │ (gRPC)  │  │ (gRPC)  │  │ (gRPC)  │         │ (gRPC)  │        │
│  │         │  │         │  │         │         │         │        │
│  │ ┌─────┐ │  │ ┌─────┐ │  │ ┌─────┐ │         │ ┌─────┐ │        │
│  │ │disk1│ │  │ │disk1│ │  │ │disk1│ │         │ │disk1│ │        │
│  │ │disk2│ │  │ │disk2│ │  │ │disk2│ │         │ │disk2│ │        │
│  │ │ ... │ │  │ │ ... │ │  │ │ ... │ │         │ │ ... │ │        │
│  │ │diskM│ │  │ │diskM│ │  │ │diskM│ │         │ │diskM│ │        │
│  │ └─────┘ │  │ └─────┘ │  │ └─────┘ │         │ └─────┘ │        │
│  └─────────┘  └─────────┘  └─────────┘         └─────────┘        │
│                                                                      │
└─────────────────────────────────────────────────────────────────────┘

2.2 Deployment Modes

Mode              MDS              DSS              Use case
────────────────  ───────────────  ───────────────  ─────────────────
Standalone        embedded in      same process     Dev / edge / IoT
                  rustfs binary                     (single node)

Small cluster     1 MDS (no HA)   3-8 DSS nodes    Small AI lab
                                                    (< 50 GPUs)

Production        3-5 MDS Raft    8-128 DSS nodes   AI training center
                  cluster                           (50-1000 GPUs)

Hyperscale        Sharded MDS     128+ DSS nodes    Cloud-scale
                  (DNE: subtree                     (1000+ GPUs,
                   partitioning)                     PB-scale)

2.3 Data Flow Summary

                    ┌──────────────────────────────────┐
                    │         CLIENT REQUEST            │
                    └────────────────┬─────────────────┘
                                     │
                    ┌────────────────┴─────────────────┐
                    │         WHAT KIND OF OP?          │
                    └──┬──────────┬──────────┬─────────┘
                       │          │          │
              metadata op    data read   data write
                       │          │          │
                       ▼          │          │
                ┌──────────┐     │          │
                │   MDS    │     │          │
                │ (lookup, │     │          │
                │  stat,   │◄────┤          │
                │  mkdir,  │     │          │
                │  create) │     │          │
                └────┬─────┘     │          │
                     │           │          │
              returns chunk      │          │
              map + lease        │          │
                     │           │          │
                     ▼           ▼          ▼
                ┌────────────────────────────────┐
                │    DSS (Data Storage Servers)   │
                │                                 │
                │  read: parallel chunk fetch     │
                │  write: parallel chunk write    │
                │         + MDS commit            │
                └─────────────────────────────────┘

3. Core Data Structures

3.1 Inode

The universal identity for every object in the unified namespace.

// crate: rustfs-inode
// file:  crates/inode/src/types.rs

/// 64-bit inode identifier.
///
/// Layout:
///   bits [63..48] = MDS partition ID (supports 65536 MDS shards)
///   bits [47..0]  = local sequence number (281 trillion inodes per shard)
///
/// Partition 0 is the default single-MDS partition.
/// ROOT_INODE = 0x0000_0000_0000_0001
pub type InodeId = u64;

pub const ROOT_INODE: InodeId = 1;

/// Inode kinds — determines which protocol views can access this inode.
#[derive(Clone, Debug, PartialEq, Eq, Serialize, Deserialize)]
pub enum InodeKind {
    RegularFile,
    Directory,
    Symlink { target: String },
    BlockVolume { volume_id: VolumeId },
}

/// Full POSIX inode attributes.
///
/// Stored in MDS RocksDB under CF "inodes", key = big-endian InodeId bytes.
/// Serialized as MessagePack for compact storage and fast decode.
#[derive(Clone, Debug, Serialize, Deserialize)]
pub struct InodeAttr {
    pub ino:        InodeId,
    pub kind:       InodeKind,
    pub size:       u64,
    pub blocks:     u64,            // in 512-byte units (for stat.st_blocks)
    pub atime:      SystemTime,
    pub mtime:      SystemTime,
    pub ctime:      SystemTime,
    pub birthtime:  SystemTime,
    pub perm:       u16,            // rwxrwxrwx (9 bits)
    pub uid:        u32,
    pub gid:        u32,
    pub nlink:      u32,
    pub rdev:       u32,
    pub flags:      u32,
    pub generation: u64,            // for NFS export and FUSE

    // Storage policy (set at creation, changeable via admin API)
    pub storage_policy: StoragePolicy,

    // S3 compatibility fields (populated when accessed via S3)
    pub content_type:     Option<String>,
    pub content_encoding: Option<String>,
    pub etag:             Option<String>,
    pub s3_metadata:      BTreeMap<String, String>,  // x-amz-meta-*
    pub s3_tags:          BTreeMap<String, String>,   // S3 object tags
    pub s3_storage_class: Option<String>,
}

/// Storage policy determines how data chunks are protected.
///
/// Selected per-file or per-directory (inherited by children).
/// AI training data → EC striping (write-once, read-many, max throughput)
/// VM block volumes → replication (random 4K writes, low latency)
/// General POSIX    → configurable (default EC for large, replica for small)
#[derive(Clone, Debug, Serialize, Deserialize)]
pub enum StoragePolicy {
    /// Reed-Solomon erasure coding across N data + M parity shards.
    /// Best for: large files, sequential I/O, write-once-read-many.
    /// Inherited from existing ecstore EC engine.
    ErasureCoded {
        data_shards:   u32,         // e.g., 4
        parity_shards: u32,         // e.g., 2
        chunk_size:    u64,         // e.g., 4 MiB
    },
    /// N-way replication.
    /// Best for: small files, random I/O, block volumes.
    Replicated {
        factor:     u32,            // e.g., 3
        chunk_size: u64,            // e.g., 4 MiB (file) or 4 MiB (block)
    },
    /// Tiered: hot data replicated, cold data EC'd.
    /// Best for: mixed workloads with lifecycle management.
    Tiered {
        hot: Box<StoragePolicy>,
        cold: Box<StoragePolicy>,
        hot_threshold_days: u32,
    },
}

impl Default for StoragePolicy {
    fn default() -> Self {
        StoragePolicy::ErasureCoded {
            data_shards: 4,
            parity_shards: 2,
            chunk_size: 4 * 1024 * 1024,  // 4 MiB
        }
    }
}

3.2 Directory Entry

// crate: rustfs-inode
// file:  crates/inode/src/dentry.rs

/// A single directory entry.
///
/// Stored in MDS RocksDB under CF "dentries".
/// Key   = big-endian parent InodeId ++ name bytes (allows prefix scan)
/// Value = DirEntry as MessagePack
#[derive(Clone, Debug, Serialize, Deserialize)]
pub struct DirEntry {
    pub name:  String,
    pub ino:   InodeId,
    pub kind:  InodeKind,       // cached from inode (avoids extra lookup)
}

/// Batch directory listing result for readdir.
#[derive(Clone, Debug, Serialize, Deserialize)]
pub struct DirListing {
    pub parent:  InodeId,
    pub entries: Vec<DirEntry>,
    pub has_more: bool,
    pub next_cursor: Option<String>,  // for paginated readdir
}

3.3 Chunk Map

// crate: rustfs-inode
// file:  crates/inode/src/chunkmap.rs

/// Globally unique chunk identifier.
///
/// Format: "{dss_node_id}-{disk_id}-{local_sequence}"
/// Example: "dss03-disk07-000000004a2f"
pub type ChunkId = String;

/// Which DSS node owns a chunk.
pub type NodeId = u32;

/// Physical location of one chunk replica or EC shard.
#[derive(Clone, Debug, Serialize, Deserialize)]
pub struct ChunkLocation {
    pub chunk_id:   ChunkId,
    pub node_id:    NodeId,
    pub disk_id:    u32,
    pub version:    u64,         // monotonic, for COW detection
    pub status:     ChunkStatus,
}

#[derive(Clone, Debug, Serialize, Deserialize)]
pub enum ChunkStatus {
    Active,
    Healing,
    Stale,
}

/// Maps a file's byte ranges to chunk storage locations.
///
/// Stored in MDS RocksDB under CF "chunkmaps".
/// Key   = big-endian InodeId
/// Value = ChunkMap as MessagePack
///
/// Design: BTreeMap indexed by chunk_index (= file_offset / chunk_size).
/// This gives O(log N) range queries for any [offset, offset+len) read/write.
///
/// For EC policy: each chunk_index maps to N+M ChunkLocations
///   (one per EC shard across different DSS nodes).
/// For replica policy: each chunk_index maps to R ChunkLocations
///   (one per replica across different DSS nodes).
#[derive(Clone, Debug, Serialize, Deserialize)]
pub struct ChunkMap {
    pub ino:          InodeId,
    pub chunk_size:   u64,
    pub policy:       StoragePolicy,
    pub chunks:       BTreeMap<u64, Vec<ChunkLocation>>,
}

impl ChunkMap {
    /// Resolve which chunks cover byte range [offset, offset+length).
    ///
    /// Returns: Vec<(chunk_index, offset_within_chunk, read_length, locations)>
    ///
    /// This is the HOT PATH for every POSIX read and block read.
    /// Must be O(log N + K) where K = number of chunks in the range.
    pub fn resolve_range(
        &self,
        offset: u64,
        length: u64,
    ) -> Vec<ChunkRangeEntry> {
        let start_idx = offset / self.chunk_size;
        let end_idx   = (offset + length).saturating_sub(1) / self.chunk_size;

        self.chunks.range(start_idx..=end_idx).map(|(idx, locs)| {
            let chunk_start = idx * self.chunk_size;
            let local_offset = if offset > chunk_start {
                offset - chunk_start
            } else { 0 };
            let local_end = ((idx + 1) * self.chunk_size).min(offset + length);
            let local_length = local_end - chunk_start - local_offset;

            ChunkRangeEntry {
                chunk_index:    *idx,
                offset_in_chunk: local_offset,
                length:         local_length,
                locations:      locs.clone(),
            }
        }).collect()
    }

    /// Total number of chunks (for capacity accounting).
    pub fn chunk_count(&self) -> u64 {
        self.chunks.len() as u64
    }
}

#[derive(Clone, Debug)]
pub struct ChunkRangeEntry {
    pub chunk_index:     u64,
    pub offset_in_chunk: u64,
    pub length:          u64,
    pub locations:       Vec<ChunkLocation>,
}

3.4 Block Volume

// crate: rustfs-inode
// file:  crates/inode/src/volume.rs

pub type VolumeId = String;   // e.g., "vol-00af39b2"
pub type SnapshotId = String;

/// Block volume descriptor.
///
/// A volume is represented in the unified namespace as:
///   InodeKind::BlockVolume { volume_id }
///   at path /volumes/{volume_name}
///
/// The volume's data is stored as an extent-mapped set of chunks,
/// reusing the same ChunkMap structure but with different access semantics:
///   - Block I/O is always aligned to volume.block_size
///   - Writes are partial (4K granularity within 4 MiB chunks)
///   - Flush / write-barrier semantics are enforced
///
/// Stored in MDS RocksDB under CF "volumes", key = VolumeId.
#[derive(Clone, Debug, Serialize, Deserialize)]
pub struct Volume {
    pub id:           VolumeId,
    pub name:         String,           // human-readable
    pub ino:          InodeId,          // inode in /volumes/
    pub capacity:     u64,              // logical size in bytes
    pub block_size:   u32,              // 512 or 4096
    pub chunk_size:   u64,              // 4 MiB default
    pub provisioning: VolumeProvisioning,
    pub state:        VolumeState,
    pub snapshots:    Vec<SnapshotId>,
    pub created_at:   SystemTime,
    pub chunk_map:    ChunkMap,         // reuses file chunk map!

    // Storage: always replicated (not EC) for random write performance
    pub replica_factor: u32,            // default 3
}

#[derive(Clone, Debug, Serialize, Deserialize)]
pub enum VolumeProvisioning {
    Thin,       // allocate on first write (default)
    Thick,      // pre-allocate all chunks at creation
}

#[derive(Clone, Debug, Serialize, Deserialize)]
pub enum VolumeState {
    Available,
    Attached { target: String },  // which host / initiator
    Snapshotting,
    Deleting,
}

/// COW snapshot of a volume.
///
/// Shares the parent's chunk map at snapshot time.
/// Subsequent writes to parent or snapshot COW the affected chunks.
#[derive(Clone, Debug, Serialize, Deserialize)]
pub struct VolumeSnapshot {
    pub id:           SnapshotId,
    pub volume_id:    VolumeId,
    pub created_at:   SystemTime,
    pub capacity:     u64,
    pub chunk_map:    ChunkMap,          // frozen copy at snapshot time
    pub cow_version:  u64,               // version counter for COW tracking
}

3.5 Leases and Locks

// crate: rustfs-inode
// file:  crates/inode/src/lease.rs

pub type ClientId = u64;
pub type SessionId = u64;

/// Client lease — enables client-side caching with coherency.
///
/// Inspired by Lustre LDLM and NFSv4 delegations.
///
/// The MDS grants leases; clients cache data/metadata locally.
/// When a conflicting operation arrives, MDS revokes the lease
/// (callback to client) before proceeding.
///
/// This is the KEY mechanism for AI training performance:
/// hundreds of readers get ReadData leases on training files →
/// each reader caches chunks locally → no MDS/DSS traffic on
/// repeated epoch reads → GPU feed rate maximized.
#[derive(Clone, Debug, Serialize, Deserialize)]
pub struct Lease {
    pub id:         u64,
    pub ino:        InodeId,
    pub client_id:  ClientId,
    pub session_id: SessionId,
    pub kind:       LeaseKind,
    pub granted_at: SystemTime,
    pub expires_at: SystemTime,       // must heartbeat to renew
}

#[derive(Clone, Debug, PartialEq, Eq, Serialize, Deserialize)]
pub enum LeaseKind {
    /// Client can cache inode attributes (stat cache).
    /// Revoked when any client modifies attributes.
    ReadMeta,

    /// Client can cache file data in its page cache.
    /// Revoked when any client writes to this file.
    /// Multiple ReadData leases can coexist.
    ReadData,

    /// Client has exclusive write access and can cache writes.
    /// Only one WriteData lease per inode.
    /// Revoked when another client requests any access.
    WriteData,

    /// Client knows the chunk map layout and can issue direct
    /// chunk reads to DSS without re-checking MDS.
    /// Revoked when file is re-striped, truncated, or extended.
    Layout,
}

/// POSIX byte-range lock (flock / fcntl).
///
/// Stored in MDS memory (not persisted — locks die with the MDS).
/// For HA MDS, locks are part of the Raft state machine.
#[derive(Clone, Debug, Serialize, Deserialize)]
pub struct ByteRangeLock {
    pub ino:       InodeId,
    pub owner:     (ClientId, u64),   // (client, pid or lock_owner)
    pub kind:      RangeLockKind,
    pub start:     u64,
    pub end:       u64,               // u64::MAX = to EOF
}

#[derive(Clone, Debug, PartialEq, Eq, Serialize, Deserialize)]
pub enum RangeLockKind {
    Read,       // F_RDLCK — shared
    Write,      // F_WRLCK — exclusive
}

3.6 MDS RocksDB Schema

Column Family          Key                            Value (msgpack)
─────────────────────  ─────────────────────────────  ───────────────────────
"inodes"               BE(inode_id)                   InodeAttr
"dentries"             BE(parent_ino) ++ name_bytes   DirEntry
"chunkmaps"            BE(inode_id)                   ChunkMap
"volumes"              volume_id_bytes                Volume
"snapshots"            snapshot_id_bytes              VolumeSnapshot
"xattrs"               BE(inode_id) ++ xattr_name     xattr_value bytes
"symlinks"             BE(inode_id)                   target_path string
"s3multipart"          upload_id_bytes                MultipartUploadState

Meta keys (default CF):
  "next_inode"         →  u64 (atomic inode counter)
  "cluster_id"         →  UUID
  "mds_version"        →  u32

(BE = big-endian encoding for correct byte-order sorting in RocksDB)

4. MDS Service Definition

4.1 gRPC API

// file: proto/mds.proto
syntax = "proto3";
package rustfs.mds;

// ═══════════════════════════════════════════════════════
//  SESSION MANAGEMENT
// ═══════════════════════════════════════════════════════

service SessionService {
    // Client registration. Returns client_id and session_id.
    // Client must call Heartbeat periodically (default: 30s).
    // If heartbeat lapses, MDS revokes all leases and locks.
    rpc OpenSession(OpenSessionReq)    returns (OpenSessionResp);
    rpc Heartbeat(HeartbeatReq)        returns (HeartbeatResp);
    rpc CloseSession(CloseSessionReq)  returns (CloseSessionResp);
}

message OpenSessionReq {
    string client_hostname = 1;
    string client_version  = 2;
    uint32 max_read_ahead_mb = 3;   // client's read-ahead capability
}

message OpenSessionResp {
    uint64 client_id  = 1;
    uint64 session_id = 2;
    uint32 heartbeat_interval_ms = 3;   // server-chosen interval
    repeated string dss_endpoints = 4;  // DSS node addresses for data path
}

// ═══════════════════════════════════════════════════════
//  NAMESPACE OPERATIONS
// ═══════════════════════════════════════════════════════

service NamespaceService {
    // Path resolution
    rpc Lookup(LookupReq)              returns (LookupResp);

    // Inode attributes
    rpc GetAttr(GetAttrReq)            returns (GetAttrResp);
    rpc SetAttr(SetAttrReq)            returns (SetAttrResp);

    // File lifecycle
    rpc Create(CreateReq)              returns (CreateResp);
    rpc Unlink(UnlinkReq)              returns (UnlinkResp);
    rpc Rename(RenameReq)              returns (RenameResp);
    rpc Link(LinkReq)                  returns (LinkResp);
    rpc Symlink(SymlinkReq)            returns (SymlinkResp);

    // Directory operations
    rpc MkDir(MkDirReq)                returns (MkDirResp);
    rpc RmDir(RmDirReq)                returns (RmDirResp);
    rpc ReadDir(ReadDirReq)            returns (ReadDirResp);

    // Extended attributes
    rpc GetXAttr(GetXAttrReq)          returns (GetXAttrResp);
    rpc SetXAttr(SetXAttrReq)          returns (SetXAttrResp);
    rpc ListXAttr(ListXAttrReq)        returns (ListXAttrResp);
    rpc RemoveXAttr(RemoveXAttrReq)    returns (RemoveXAttrResp);
}

message LookupReq {
    uint64 session_id   = 1;
    uint64 parent_ino   = 2;
    string name         = 3;
}

message LookupResp {
    uint64    ino        = 1;
    InodeAttr attr       = 2;
    uint64    generation = 3;
}

message CreateReq {
    uint64 session_id    = 1;
    uint64 parent_ino    = 2;
    string name          = 3;
    uint32 mode          = 4;    // permission bits
    uint32 uid           = 5;
    uint32 gid           = 6;
    uint32 flags         = 7;    // O_EXCL, O_TRUNC, etc.
    StoragePolicy policy = 8;    // optional override
}

message CreateResp {
    uint64    ino        = 1;
    InodeAttr attr       = 2;
    uint64    generation = 3;
    Lease     lease      = 4;    // auto-granted write lease
}

message ReadDirReq {
    uint64 session_id = 1;
    uint64 ino        = 2;
    uint64 offset     = 3;     // cookie for pagination
    uint32 limit      = 4;     // max entries to return
    bool   plus       = 5;     // readdirplus: include attrs
}

message ReadDirResp {
    repeated DirEntryProto entries = 1;
    bool   has_more   = 2;
    uint64 next_offset = 3;
}

// ═══════════════════════════════════════════════════════
//  DATA LAYOUT (CHUNK MAP) OPERATIONS
// ═══════════════════════════════════════════════════════

service DataLayoutService {
    // Get chunk map for a file (used before data I/O)
    rpc GetChunkMap(GetChunkMapReq)    returns (GetChunkMapResp);

    // Allocate new chunks for writes (MDS picks DSS nodes)
    rpc AllocChunks(AllocChunksReq)    returns (AllocChunksResp);

    // Commit chunk writes (update chunk map after successful DSS writes)
    rpc CommitWrite(CommitWriteReq)    returns (CommitWriteResp);

    // Truncate file to new size (free chunks beyond new size)
    rpc Truncate(TruncateReq)          returns (TruncateResp);
}

message GetChunkMapReq {
    uint64 session_id = 1;
    uint64 ino        = 2;
    uint64 offset     = 3;   // start of range (0 = entire file)
    uint64 length     = 4;   // 0 = to EOF
}

message GetChunkMapResp {
    ChunkMapProto chunk_map = 1;
    Lease         layout_lease = 2;   // optional layout lease
}

message AllocChunksReq {
    uint64 session_id    = 1;
    uint64 ino           = 2;
    uint64 start_index   = 3;   // first chunk index to allocate
    uint32 count         = 4;   // number of chunks to allocate
    StoragePolicy policy = 5;
}

message AllocChunksResp {
    repeated ChunkAllocation allocations = 1;
}

message ChunkAllocation {
    uint64 chunk_index = 1;
    repeated ChunkLocationProto locations = 2;  // where to write
}

message CommitWriteReq {
    uint64 session_id = 1;
    uint64 ino        = 2;
    uint64 new_size   = 3;     // file size after this write
    int64  mtime_ns   = 4;     // new mtime
    repeated ChunkCommit chunks = 5;
}

message ChunkCommit {
    uint64 chunk_index = 1;
    repeated ChunkLocationProto locations = 2;
    uint64 bytes_written = 3;
}

// ═══════════════════════════════════════════════════════
//  LEASE AND LOCK OPERATIONS
// ═══════════════════════════════════════════════════════

service LeaseService {
    rpc AcquireLease(AcquireLeaseReq)  returns (AcquireLeaseResp);
    rpc RenewLease(RenewLeaseReq)      returns (RenewLeaseResp);
    rpc ReleaseLease(ReleaseLeaseReq)  returns (ReleaseLeaseResp);

    // Server → Client callback for lease revocation.
    // Client implements this as a gRPC server endpoint.
    // MDS calls this when a conflicting operation requires revocation.
    rpc RevokeLease(RevokeLeaseReq)    returns (RevokeLeaseResp);

    // POSIX byte-range locks
    rpc LockRange(LockRangeReq)        returns (LockRangeResp);
    rpc UnlockRange(UnlockRangeReq)    returns (UnlockRangeResp);
    rpc TestLock(TestLockReq)          returns (TestLockResp);
}

// ═══════════════════════════════════════════════════════
//  BLOCK VOLUME OPERATIONS
// ═══════════════════════════════════════════════════════

service VolumeService {
    rpc CreateVolume(CreateVolumeReq)     returns (CreateVolumeResp);
    rpc DeleteVolume(DeleteVolumeReq)     returns (DeleteVolumeResp);
    rpc ResizeVolume(ResizeVolumeReq)     returns (ResizeVolumeResp);
    rpc AttachVolume(AttachVolumeReq)     returns (AttachVolumeResp);
    rpc DetachVolume(DetachVolumeReq)     returns (DetachVolumeResp);
    rpc SnapshotVolume(SnapshotReq)       returns (SnapshotResp);
    rpc CloneVolume(CloneVolumeReq)       returns (CloneVolumeResp);
    rpc ListVolumes(ListVolumesReq)       returns (ListVolumesResp);
    rpc GetVolumeChunkMap(GetVolMapReq)   returns (GetVolMapResp);
}

message CreateVolumeReq {
    string name           = 1;
    uint64 capacity_bytes = 2;
    uint32 block_size     = 3;   // 512 or 4096
    string provisioning   = 4;   // "thin" or "thick"
    uint32 replica_factor = 5;   // default 3
}

4.2 S3 API Adaptation

The existing rustfs S3 API server calls into MDS instead of directly into ecstore for namespace operations:

// file: rustfs/src/storage/unified.rs (NEW — replaces parts of ecfs.rs)

/// Unified storage adapter.
///
/// Translates S3 API calls into MDS + DSS operations.
/// Installed as the ObjectLayer implementation when unified mode is enabled.
pub struct UnifiedObjectLayer {
    mds:         MdsClient,
    dss_pool:    DssConnectionPool,
    s3_compat:   S3CompatLayer,       // handles S3-specific quirks
}

impl ObjectLayer for UnifiedObjectLayer {
    async fn put_object(&self, bucket: &str, key: &str, data: Stream, opts: PutOpts)
        -> Result<ObjectInfo>
    {
        // 1. Resolve path: /buckets/{bucket}/{key}
        let parent_ino = self.resolve_s3_path(bucket, key).await?;

        // 2. Create inode via MDS (or overwrite existing)
        let create_resp = self.mds.create(CreateReq {
            parent_ino,
            name: key_last_component(key),
            mode: 0o644,
            policy: self.policy_for_bucket(bucket),
            ..Default::default()
        }).await?;

        // 3. Allocate chunks via MDS
        let chunk_count = (data.content_length + chunk_size - 1) / chunk_size;
        let alloc = self.mds.alloc_chunks(AllocChunksReq {
            ino: create_resp.ino,
            count: chunk_count as u32,
            ..Default::default()
        }).await?;

        // 4. Write data to DSS nodes (parallel)
        let commits = self.write_chunks_parallel(&alloc, data).await?;

        // 5. Commit to MDS
        self.mds.commit_write(CommitWriteReq {
            ino: create_resp.ino,
            new_size: data.content_length,
            chunks: commits,
            ..Default::default()
        }).await?;

        // 6. Return S3-compatible ObjectInfo
        Ok(self.s3_compat.to_object_info(&create_resp.attr))
    }

    async fn get_object(&self, bucket: &str, key: &str, range: Option<HttpRange>)
        -> Result<GetObjectResp>
    {
        // 1. Resolve path → inode
        let ino = self.resolve_s3_key(bucket, key).await?;

        // 2. Get chunk map from MDS
        let (offset, length) = range.unwrap_or((0, u64::MAX));
        let cm = self.mds.get_chunk_map(ino, offset, length).await?;

        // 3. Read chunks from DSS (parallel — reuses existing ecstore logic)
        let stream = self.read_chunks_parallel(&cm, offset, length).await?;

        Ok(GetObjectResp { stream, attr: cm.attr })
    }

    async fn list_objects(&self, bucket: &str, prefix: &str, opts: ListOpts)
        -> Result<ListObjectsResp>
    {
        // Translate to MDS readdir on the resolved directory.
        // If prefix contains "/", walk directory tree.
        // Much faster than scanning all keys!
        let dir_ino = self.resolve_s3_prefix(bucket, prefix).await?;
        let listing = self.mds.read_dir(dir_ino, opts.marker, opts.max_keys).await?;
        Ok(self.s3_compat.to_list_response(listing, prefix, opts.delimiter))
    }
}

5. DSS Chunk Service

5.1 gRPC API

// file: proto/chunk.proto
syntax = "proto3";
package rustfs.chunk;

service ChunkService {
    // Write data to a chunk at an offset.
    //
    // For replicated chunks: direct byte-range write.
    // For EC chunks: if write covers full EC block(s), encode + write shards.
    //               if partial, read-modify-write the affected block.
    rpc WriteChunk(stream WriteChunkReq)  returns (WriteChunkResp);

    // Read data from a chunk at an offset.
    //
    // This reuses and extends the existing rustfs RPC:
    //   /rustfs/rpc/read_file_stream?disk=...&offset=...&length=...
    rpc ReadChunk(ReadChunkReq)           returns (stream ReadChunkResp);

    // Batch read: read multiple chunk ranges in one RPC.
    // Reduces round trips for striped reads.
    rpc ReadChunkBatch(ReadChunkBatchReq) returns (stream ReadChunkBatchResp);

    // Ensure chunk data is durable (fsync).
    rpc FlushChunk(FlushChunkReq)         returns (FlushChunkResp);

    // Truncate or delete chunk.
    rpc TruncateChunk(TruncateChunkReq)   returns (TruncateChunkResp);
    rpc DeleteChunks(DeleteChunksReq)      returns (DeleteChunksResp);

    // Status
    rpc StatChunk(StatChunkReq)            returns (StatChunkResp);
}

message WriteChunkReq {
    string chunk_id    = 1;
    uint64 offset      = 2;    // byte offset within chunk
    bytes  data        = 3;    // payload (streamed for large writes)
    bool   sync        = 4;    // fsync after write
    bool   create      = 5;    // create chunk if not exists
}

message WriteChunkResp {
    uint64 bytes_written = 1;
    string checksum      = 2;   // new checksum after write
}

message ReadChunkReq {
    string chunk_id  = 1;
    uint64 offset    = 2;
    uint64 length    = 3;
    bool   direct_io = 4;      // bypass OS page cache
    bool   verify    = 5;      // verify bitrot checksum
}

message ReadChunkResp {
    bytes  data   = 1;         // streamed
    uint64 offset = 2;         // offset of this chunk within stream
}

message ReadChunkBatchReq {
    repeated ReadChunkReq reads = 1;
}

message ReadChunkBatchResp {
    uint32 index = 1;          // which request in the batch
    bytes  data  = 2;          // streamed
}

5.2 DSS Chunk Storage on Disk

Physical disk layout per DSS node:

/{volume_mount}/
├── .rustfs.sys/                     # cluster metadata (existing)
├── .chunks/                         # chunk data (NEW)
│   ├── 00/                          # first 2 hex chars of chunk_id
│   │   ├── 00a3f9b2.dat            # chunk data file
│   │   ├── 00a3f9b2.meta           # chunk metadata (checksum, size)
│   │   └── 00b1cc07.dat
│   ├── 01/
│   │   └── ...
│   └── ff/
│       └── ...
├── {bucket}/                        # S3 legacy object data (existing)
│   └── {key_hash}/
│       ├── xl.meta
│       └── part.1
└── .journal/                        # write-ahead log for block volumes
    ├── wal-000001.log
    └── wal-000002.log

5.3 Chunk Write Strategies

// crate: rustfs-chunk-service
// file:  crates/chunk-service/src/write.rs

impl ChunkServiceImpl {
    /// Route write to the correct strategy based on storage policy.
    pub async fn write_chunk(
        &self,
        req: WriteChunkReq,
        policy: &StoragePolicy,
    ) -> Result<WriteChunkResp> {
        match policy {
            StoragePolicy::Replicated { .. } => {
                self.write_replicated(req).await
            }
            StoragePolicy::ErasureCoded { data_shards, parity_shards, .. } => {
                let block_size = EC_BLOCK_SIZE; // 1 MiB
                let write_size = req.data.len() as u64;

                if write_size >= block_size
                    && req.offset % block_size == 0
                {
                    // Full-block write: fast path, no RMW
                    self.write_ec_full_blocks(req, *data_shards, *parity_shards).await
                } else {
                    // Partial write: read-modify-write
                    self.write_ec_partial(req, *data_shards, *parity_shards).await
                }
            }
            _ => unreachable!(),
        }
    }

    /// Replicated write: write to all replicas in parallel.
    ///
    /// Latency = max(replica_write_latencies) — bounded by slowest disk.
    /// For block volumes, this is the primary write path.
    async fn write_replicated(&self, req: WriteChunkReq) -> Result<WriteChunkResp> {
        let chunk_path = self.chunk_path(&req.chunk_id);

        // Open chunk file (create if necessary)
        let file = tokio::fs::OpenOptions::new()
            .write(true)
            .create(req.create)
            .open(&chunk_path)
            .await?;

        // Write at offset
        file.seek(SeekFrom::Start(req.offset)).await?;
        file.write_all(&req.data).await?;

        if req.sync {
            file.sync_data().await?;
        }

        // Update bitrot checksum for the affected range
        let checksum = self.update_chunk_checksum(
            &req.chunk_id, req.offset, &req.data
        ).await?;

        Ok(WriteChunkResp {
            bytes_written: req.data.len() as u64,
            checksum,
        })
    }

    /// EC full-block write: encode data blocks → parity blocks, write all shards.
    ///
    /// Reuses existing ecstore erasure_coding::encode path.
    async fn write_ec_full_blocks(
        &self,
        req: WriteChunkReq,
        data_shards: u32,
        parity_shards: u32,
    ) -> Result<WriteChunkResp> {
        // Split into EC block-aligned segments
        // Encode each block → data + parity shards
        // Write each shard to its designated disk
        // (This is the existing ecstore SetDisks::put_object logic,
        //  refactored to operate on individual blocks rather than whole objects)
        todo!("Reuse ecstore erasure_coding::encode")
    }
}

6. Client Implementations

6.1 FUSE Client (rustfs-mount)

Binary: rustfs-mount
Usage:  rustfs-mount --mds=mds1:9100,mds2:9100 --mountpoint=/mnt/rustfs [options]

Options:
  --read-ahead=32m         Max read-ahead window per file
  --cache-size=4g          Client-side data cache (page cache supplement)
  --attr-timeout=5s        Attribute cache TTL (overridden by leases)
  --entry-timeout=5s       Dentry cache TTL (overridden by leases)
  --direct-io              Default to O_DIRECT (bypass client cache)
  --storage-policy=ec:4+2  Default storage policy for new files
  --worker-threads=16      FUSE worker threads
  --max-write=1m           Max write size per FUSE request

6.1.1 FUSE Client Architecture

// crate: rustfs-fuse
// file:  crates/fuse-client/src/client.rs

pub struct RustfsClient {
    // Connections
    mds:            MdsClient,
    dss_pool:       DssConnectionPool,

    // Caches
    inode_cache:    LruCache<InodeId, CachedInode>,
    dentry_cache:   LruCache<(InodeId, String), CachedDentry>,
    chunk_map_cache: LruCache<InodeId, Arc<ChunkMap>>,

    // State
    open_files:     DashMap<FileHandle, OpenFile>,
    session:        Session,
    next_fh:        AtomicU64,

    // Engines
    read_ahead:     ReadAheadEngine,
    lease_manager:  ClientLeaseManager,
    write_buffer:   WriteBufferManager,
}

struct OpenFile {
    ino:        InodeId,
    flags:      u32,
    chunk_map:  Arc<ChunkMap>,
    lease:      Option<Lease>,
    dirty:      AtomicBool,
}

struct CachedInode {
    attr:       InodeAttr,
    valid_until: Instant,
    lease:      Option<Lease>,       // if lease held, cache is valid until revoked
}

6.1.2 FUSE Read Path (Hot Path for AI)

// crate: rustfs-fuse
// file:  crates/fuse-client/src/read.rs

impl RustfsClient {
    /// FUSE read handler. Called for every read() syscall.
    ///
    /// Performance target: <5ms first-byte latency for cached chunk maps,
    /// aggregate bandwidth scales linearly with DSS node count.
    pub async fn fuse_read(
        &self,
        ino: InodeId,
        fh: FileHandle,
        offset: u64,
        size: u32,
    ) -> Result<Vec<u8>> {
        let file = self.open_files.get(&fh)
            .ok_or(Error::BadFileHandle)?;

        // Step 1: Get chunk map (cached if layout lease held)
        let chunk_map = self.get_chunk_map_cached(ino).await?;

        // Step 2: Resolve byte range → chunk reads
        let chunk_reads = chunk_map.resolve_range(offset, size as u64);

        // Step 3: Issue parallel reads to DSS nodes
        //
        // KEY OPTIMIZATION: each chunk maps to a different DSS node
        // (due to striping), so all reads go in parallel to different
        // servers. For a 64 MiB read with 4 MiB chunks = 16 parallel
        // fetches to up to 16 different DSS nodes.
        let mut fetches = Vec::with_capacity(chunk_reads.len());
        for entry in &chunk_reads {
            let loc = self.pick_best_location(&entry.locations);
            let dss = self.dss_pool.get(loc.node_id);
            fetches.push(dss.read_chunk(ReadChunkReq {
                chunk_id:  loc.chunk_id.clone(),
                offset:    entry.offset_in_chunk,
                length:    entry.length,
                direct_io: file.flags & O_DIRECT != 0,
                verify:    true,
            }));
        }
        let results = futures::future::try_join_all(fetches).await?;

        // Step 4: Assemble results into contiguous buffer
        let mut buf = Vec::with_capacity(size as usize);
        for result in results {
            buf.extend_from_slice(&result.data);
        }

        // Step 5: Trigger read-ahead
        self.read_ahead.on_read(ino, offset, size as u64, &chunk_map);

        Ok(buf)
    }

    /// Pick the best replica/shard location for a read.
    ///
    /// Priority: local node > same rack > lowest latency > round-robin
    fn pick_best_location(&self, locs: &[ChunkLocation]) -> &ChunkLocation {
        // 1. Check if any location is on this client's local node
        //    (for DSS nodes that are also compute nodes)
        if let Some(local) = locs.iter().find(|l| l.node_id == self.local_node_id) {
            return local;
        }
        // 2. Pick by lowest recent latency
        locs.iter()
            .min_by_key(|l| self.dss_pool.latency_ns(l.node_id))
            .unwrap()
    }
}

6.1.3 Read-Ahead Engine

// crate: rustfs-fuse
// file:  crates/fuse-client/src/readahead.rs

/// Adaptive read-ahead engine.
///
/// Detects sequential, strided, and random access patterns per-file.
/// For sequential access, exponentially grows prefetch window up to max.
/// For strided access (common in DataLoader: read shard, jump, read next),
/// learns the stride and prefetches accordingly.
///
/// Design: mirrors Linux kernel's readahead algorithm but runs in userspace
/// and operates at chunk granularity rather than page granularity.
pub struct ReadAheadEngine {
    trackers:   DashMap<InodeId, AccessTracker>,
    prefetcher: mpsc::Sender<PrefetchCmd>,
    config:     ReadAheadConfig,
}

pub struct ReadAheadConfig {
    pub max_window:       u64,    // bytes (default: 32 MiB)
    pub initial_window:   u64,    // bytes (default: chunk_size)
    pub growth_factor:    f64,    // default: 2.0
    pub shrink_threshold: u32,    // sequential break count to shrink (default: 2)
}

struct AccessTracker {
    last_offset:   u64,
    last_size:     u64,
    window:        u64,
    sequential_n:  u32,
    stride:        Option<u64>,
    stride_hits:   u32,
    pattern:       AccessPattern,
}

enum AccessPattern {
    Unknown,
    Sequential,
    Strided { stride: u64 },
    Random,
}

impl ReadAheadEngine {
    /// Called after every successful read.
    pub fn on_read(
        &self,
        ino: InodeId,
        offset: u64,
        size: u64,
        chunk_map: &ChunkMap,
    ) {
        let mut tracker = self.trackers.entry(ino).or_default();

        // Classify access pattern
        let expected_next = tracker.last_offset + tracker.last_size;
        let is_sequential = offset == expected_next;
        let is_strided = tracker.stride.map_or(false, |s| {
            offset == tracker.last_offset + s
        });

        match (is_sequential, is_strided) {
            (true, _) => {
                tracker.pattern = AccessPattern::Sequential;
                tracker.sequential_n += 1;
                tracker.window = (tracker.window as f64 * self.config.growth_factor)
                    as u64;
                tracker.window = tracker.window.min(self.config.max_window);
            }
            (false, true) => {
                tracker.pattern = AccessPattern::Strided {
                    stride: tracker.stride.unwrap()
                };
                tracker.stride_hits += 1;
                tracker.window = (tracker.window as f64 * self.config.growth_factor)
                    as u64;
                tracker.window = tracker.window.min(self.config.max_window);
            }
            (false, false) => {
                // Try to learn a new stride
                let gap = offset.wrapping_sub(tracker.last_offset);
                if gap > 0 && gap < chunk_map.chunk_size * 1024 {
                    tracker.stride = Some(gap);
                    tracker.stride_hits = 0;
                }
                tracker.sequential_n = 0;
                tracker.window = self.config.initial_window;
            }
        }

        // Issue prefetch
        if tracker.window > 0 {
            let prefetch_start = match tracker.pattern {
                AccessPattern::Sequential => offset + size,
                AccessPattern::Strided { stride } => offset + stride,
                _ => return,
            };

            let chunks = chunk_map.resolve_range(prefetch_start, tracker.window);
            let _ = self.prefetcher.try_send(PrefetchCmd {
                ino,
                chunks,
                priority: tracker.sequential_n,
            });
        }

        tracker.last_offset = offset;
        tracker.last_size = size;
    }
}

6.2 Block Device Exports

6.2.1 iSCSI Target

// crate: rustfs-iscsi
// file:  crates/iscsi-target/src/target.rs

/// iSCSI target that exports RustFS block volumes as SCSI LUNs.
///
/// Each volume maps to one LUN on the target.
/// Supports multi-path: volume can be exported from multiple targets.
///
/// Target naming: iqn.2026-06.com.rustfs:{volume_name}
pub struct RustfsIscsiTarget {
    mds:       MdsClient,
    dss_pool:  DssConnectionPool,
    volumes:   DashMap<Lun, AttachedVolume>,
    wal:       WriteAheadLog,
}

struct AttachedVolume {
    volume:    Volume,
    chunk_map: ChunkMap,        // cached from MDS
    dirty_map: RwLock<BTreeSet<u64>>,  // dirty chunk indices
}

impl ScsiTarget for RustfsIscsiTarget {
    async fn execute_command(&self, cmd: ScsiCommand) -> ScsiResult {
        match cmd.opcode {
            SCSI_READ_16 => {
                let lba = cmd.lba();
                let blocks = cmd.transfer_length();
                let vol = self.volumes.get(&cmd.lun)?;

                let byte_offset = lba * vol.volume.block_size as u64;
                let byte_length = blocks as u64 * vol.volume.block_size as u64;

                // Resolve chunks and parallel read
                let data = self.read_volume_range(
                    &vol, byte_offset, byte_length
                ).await?;

                ScsiResult::data(data)
            }

            SCSI_WRITE_16 => {
                let lba = cmd.lba();
                let vol = self.volumes.get(&cmd.lun)?;
                let byte_offset = lba * vol.volume.block_size as u64;

                // WAL entry for write ordering
                self.wal.append(WalEntry::Write {
                    volume_id: vol.volume.id.clone(),
                    offset: byte_offset,
                    data: cmd.data().to_vec(),
                }).await?;

                // Write to chunk storage
                self.write_volume_range(
                    &vol, byte_offset, cmd.data()
                ).await?;

                ScsiResult::ok()
            }

            SCSI_SYNCHRONIZE_CACHE => {
                let vol = self.volumes.get(&cmd.lun)?;
                self.flush_volume(&vol).await?;
                ScsiResult::ok()
            }

            SCSI_INQUIRY => {
                ScsiResult::inquiry(InquiryData {
                    vendor: b"RustFS  ",
                    product: b"Block Volume    ",
                    revision: b"0100",
                })
            }

            SCSI_READ_CAPACITY_16 => {
                let vol = self.volumes.get(&cmd.lun)?;
                let blocks = vol.volume.capacity / vol.volume.block_size as u64;
                ScsiResult::capacity(blocks - 1, vol.volume.block_size)
            }

            _ => ScsiResult::check_condition(ILLEGAL_REQUEST),
        }
    }
}

6.2.2 NVMe-oF Target (High-Performance Path)

Binary: rustfs-nvmeof
Usage:  rustfs-nvmeof --mds=mds1:9100 --transport=tcp --port=4420

Exposes volumes as NVMe namespaces.
For RDMA transport, use --transport=rdma (requires RDMA NIC).

Implementation: wraps SPDK NVMe-oF target library via FFI,
  or uses kernel nvmet for TCP transport.

Maps:
  NVMe Read  → ReadChunk (parallel per chunk)
  NVMe Write → WriteChunk (replicated)
  NVMe Flush → FlushChunk
  NVMe TRIM  → TruncateChunk (deallocate)

7. Performance Design

7.1 AI Read Bandwidth Target

Goal: saturate GPU cluster's storage bandwidth

Cluster: 128 GPU nodes, each with 100 Gbps NIC
Aggregate demand: 128 × 12.5 GB/s = 1.6 TB/s

DSS cluster: 32 nodes, each with 4 × NVMe (3 GB/s each) + 100 Gbps NIC
Aggregate supply: 32 × 12 GB/s = 384 GB/s (disk-limited)
                  32 × 12.5 GB/s = 400 GB/s (network-limited)
Effective: ~380 GB/s

Gap: 1.6 TB/s demand vs 380 GB/s supply → 4.2x shortfall

Solution: client-side caching + read-ahead

Training epochs re-read the same data. After epoch 1:
  - 90%+ reads served from client page cache (ReadData lease held)
  - Only epoch 1 needs DSS bandwidth
  - Client needs enough RAM to cache its shard of the dataset

With 128 GPU nodes, each caching its 1/128 fraction:
  Dataset 10 TB → each node caches ~80 GB → fits in 128-256 GB RAM
  Epoch 1: 10 TB / 380 GB/s = ~26 seconds
  Epoch 2+: served from client cache → limited by memory bandwidth (~100+ GB/s per node)

7.2 Latency Targets

Operation                     Target       Notes
─────────────────────────────  ──────────  ──────────────────────────
POSIX stat (cached)           < 0.01 ms   Served from client inode cache
POSIX stat (uncached)         < 1 ms      1 MDS gRPC round trip
POSIX open                    < 2 ms      MDS lookup + chunk map fetch
POSIX read (cached chunk)     < 0.1 ms    Client page cache hit
POSIX read (1 chunk, uncached)< 5 ms      1 DSS gRPC round trip
POSIX readdir (100 entries)   < 2 ms      1 MDS gRPC round trip
POSIX create                  < 3 ms      MDS create + lease grant

Block read 4K (cached)        < 0.1 ms    iSCSI initiator cache
Block read 4K (uncached)      < 2 ms      1 DSS round trip
Block write 4K                < 1 ms      Write to nearest replica + WAL
Block flush                   < 5 ms      Sync all dirty chunks

S3 GET (small object)         < 10 ms     MDS lookup + 1 DSS read
S3 PUT (small object)         < 15 ms     MDS create + 1 DSS write + commit
S3 LIST (1000 objects)        < 5 ms      MDS readdir (O(1) not O(N)!)

7.3 Client-Side Caching Hierarchy

┌─────────────────────────────────────────────────┐
│                FUSE CLIENT CACHES                │
│                                                  │
│  ┌──────────────────────────────────────────┐   │
│  │  Level 1: Linux VFS Page Cache           │   │
│  │  (kernel manages, FUSE honors via TTL)   │   │
│  │  Size: controlled by OS (typically 50%   │   │
│  │        of free RAM)                      │   │
│  │  Eviction: LRU                           │   │
│  │  Coherency: invalidated on lease revoke  │   │
│  └──────────────────────────────────────────┘   │
│                                                  │
│  ┌──────────────────────────────────────────┐   │
│  │  Level 2: Client Chunk Cache             │   │
│  │  (userspace, managed by rustfs-mount)    │   │
│  │  Size: --cache-size (default 4 GB)       │   │
│  │  Eviction: LRU with frequency boost      │   │
│  │  Coherency: tied to ReadData lease       │   │
│  └──────────────────────────────────────────┘   │
│                                                  │
│  ┌──────────────────────────────────────────┐   │
│  │  Level 3: Inode + Dentry + ChunkMap Cache│   │
│  │  (metadata caches in rustfs-mount)       │   │
│  │  Size: 100K entries (inode), 500K (dentry│)  │
│  │  Coherency: tied to ReadMeta/Layout lease│   │
│  └──────────────────────────────────────────┘   │
│                                                  │
└─────────────────────────────────────────────────┘

8. Crate Map and Build Targets

8.1 New Crates

crates/
├── inode/                  # 3K lines — core types (InodeAttr, ChunkMap, Volume, Lease)
│   ├── src/
│   │   ├── types.rs        # InodeId, InodeAttr, InodeKind, StoragePolicy
│   │   ├── dentry.rs       # DirEntry, DirListing
│   │   ├── chunkmap.rs     # ChunkMap, ChunkLocation, resolve_range()
│   │   ├── volume.rs       # Volume, VolumeSnapshot, ExtentMap
│   │   ├── lease.rs        # Lease, LeaseKind, ByteRangeLock
│   │   └── lib.rs
│   └── Cargo.toml          # deps: serde, rmp-serde
│
├── mds/                    # 18K lines — metadata server
│   ├── src/
│   │   ├── service.rs      # gRPC service implementation (NamespaceService, etc.)
│   │   ├── store.rs        # MetadataStore trait
│   │   ├── rocksdb.rs      # RocksDB backend
│   │   ├── raft.rs         # Raft consensus wrapper (openraft)
│   │   ├── lease_mgr.rs    # Server-side lease manager + revocation
│   │   ├── lock_mgr.rs     # Byte-range lock manager
│   │   ├── alloc.rs        # Chunk allocation policy (balancing across DSS nodes)
│   │   ├── s3bridge.rs     # S3 path ↔ inode translation
│   │   ├── gc.rs           # Garbage collection for orphaned chunks
│   │   └── main.rs         # MDS binary entry point
│   └── Cargo.toml          # deps: tonic, rocksdb, openraft, rustfs-inode
│
├── chunk-service/          # 6K lines — DSS chunk I/O service
│   ├── src/
│   │   ├── service.rs      # gRPC ChunkService implementation
│   │   ├── write.rs        # Write strategies (replicated, EC full, EC partial)
│   │   ├── read.rs         # Read with bitrot verify
│   │   ├── layout.rs       # On-disk chunk file management
│   │   └── lib.rs
│   └── Cargo.toml          # deps: tonic, rustfs-io-core, rustfs-ecstore-core
│
├── fuse-client/            # 12K lines — FUSE client
│   ├── src/
│   │   ├── client.rs       # RustfsClient struct
│   │   ├── fuse_ops.rs     # fuser::Filesystem implementation
│   │   ├── read.rs         # Read path with parallel chunk fetch
│   │   ├── write.rs        # Write path with buffering + commit
│   │   ├── readahead.rs    # Adaptive read-ahead engine
│   │   ├── cache.rs        # Client-side chunk and metadata caches
│   │   ├── lease.rs        # Client-side lease manager + revocation handler
│   │   └── main.rs         # rustfs-mount binary entry point
│   └── Cargo.toml          # deps: fuser, tonic, rustfs-inode
│
├── iscsi-target/           # 8K lines — iSCSI target
│   ├── src/
│   │   ├── target.rs       # SCSI command handler
│   │   ├── session.rs      # iSCSI session/connection management
│   │   ├── wal.rs          # Write-ahead log for block write ordering
│   │   └── main.rs         # rustfs-iscsi binary entry point
│   └── Cargo.toml
│
└── nvmeof-target/          # 5K lines — NVMe-oF target (phase 2)
    ├── src/
    │   ├── target.rs       # NVMe command handler
    │   └── main.rs
    └── Cargo.toml

8.2 Modified Existing Crates

crates/ecstore/             # DECOMPOSE into:
  → ecstore-core/           # EC encode/decode, quorum, shard math
  → ecstore-disk/           # LocalDisk, disk health
  → ecstore-pool/           # EndpointServerPools, set management
  → ecstore-s3/             # S3 object operations (existing SetDisks logic)

crates/lock/                # EXTEND:
  + posix_lock.rs           # ByteRangeLock support
  + lease.rs                # Server-side lease tracking (used by MDS)

crates/filemeta/            # EXTEND:
  + posix_inode_id field in ObjectInfo (bridge to unified namespace)

crates/protos/              # ADD:
  + mds.proto               # MDS service definitions
  + chunk.proto             # DSS chunk service definitions

rustfs/src/storage/         # ADD:
  + unified.rs              # UnifiedObjectLayer (S3 API → MDS + DSS)

8.3 Build Targets

Binary               Crate Entry          Description
───────────────────  ────────────────────  ─────────────────────────────
rustfs               rustfs/src/main.rs   S3 API + admin console (existing)
                                          + unified namespace mode (new)
rustfs-mds           crates/mds/main.rs   Metadata server
rustfs-mount         crates/fuse-client/  FUSE client (POSIX mount)
                     main.rs
rustfs-iscsi         crates/iscsi-target/ iSCSI target for block volumes
                     main.rs
rustfs-nvmeof        crates/nvmeof-       NVMe-oF target (future)
                     target/main.rs

9. Configuration

9.1 Cluster Configuration

# /etc/rustfs/cluster.toml

[cluster]
id = "prod-ai-cluster-01"
mode = "unified"                # "object-only" | "unified"

[mds]
endpoints = ["mds1:9100", "mds2:9100", "mds3:9100"]
replication = "raft"            # "none" | "raft"

[dss]
# DSS nodes auto-register with MDS at startup.
# This section defines defaults.
chunk_size = "4MiB"
default_storage_policy = "ec:4+2"

[namespace]
# S3 buckets appear under this path in the POSIX namespace
bucket_root = "/buckets"
# Block volumes appear under this path
volume_root = "/volumes"
# Default path for POSIX files not associated with an S3 bucket
default_root = "/"

[leases]
heartbeat_interval = "30s"
lease_timeout = "90s"           # revoke if no heartbeat for this long
read_lease_ttl = "300s"         # max lease duration before forced renewal

[readahead]
max_window = "32MiB"
initial_window = "4MiB"

[block]
default_replica_factor = 3
wal_sync_mode = "fdatasync"     # "fdatasync" | "fsync" | "none"
thin_provision = true

9.2 Storage Policy Profiles

# /etc/rustfs/policies.toml

[profiles.ai-training]
# Optimized for write-once, read-many from hundreds of GPU nodes.
storage_policy = "ec:8+4"          # high durability, max throughput
chunk_size = "4MiB"                # matches typical DataLoader read size
default_lease = "ReadData"         # auto-grant read leases
readahead_max = "64MiB"           # aggressive prefetch

[profiles.ai-checkpoints]
# Optimized for periodic large sequential writes + occasional reads.
storage_policy = "ec:4+2"
chunk_size = "4MiB"
default_lease = "WriteData"

[profiles.vm-block]
# Optimized for random 4K I/O.
storage_policy = "replicated:3"
chunk_size = "4MiB"
default_lease = "none"             # block volumes manage coherency via iSCSI
wal_enabled = true

[profiles.scratch]
# General purpose POSIX workspace.
storage_policy = "replicated:2"    # lower durability, faster writes
chunk_size = "1MiB"

10. Implementation Phases

Phase 1 — Foundation (Months 1-3)

Deliverables:
  ✓ crates/inode/        — all types compiled and tested
  ✓ crates/mds/          — single-node RocksDB MDS with namespace + chunk map APIs
  ✓ crates/chunk-service/ — replicated chunk read/write on DSS nodes
  ✓ crates/fuse-client/  — basic FUSE mount with open/read/write/stat/readdir/mkdir

Milestone: mount /mnt/rustfs, create files, read them back, run ls -la

Acceptance test:
  $ rustfs-mds --standalone &
  $ rustfs-mount --mds=localhost:9100 /mnt/rustfs &
  $ dd if=/dev/urandom of=/mnt/rustfs/test.bin bs=4M count=256  # 1 GB write
  $ md5sum /mnt/rustfs/test.bin                                  # read back
  $ ls -la /mnt/rustfs/                                          # readdir

Phase 2 — AI-Ready (Months 3-5)

Deliverables:
  ✓ Read-ahead engine with sequential + stride detection
  ✓ ReadData lease for multi-reader caching
  ✓ Layout lease for chunk map caching
  ✓ EC-striped chunks (reuse ecstore encode/decode)
  ✓ Parallel multi-DSS chunk fetch in FUSE read path
  ✓ S3 API routed through MDS (UnifiedObjectLayer)

Milestone: train a real model using POSIX mount

Acceptance test:
  $ torchrun --nproc_per_node=8 train.py \
      --data_path=/mnt/rustfs/buckets/training-data/imagenet/
  # 8 GPU processes reading concurrently via FUSE
  # Verify: epoch 2 is faster than epoch 1 (cache warm)
  # Verify: S3 PUT to "training-data" bucket visible in /mnt/rustfs/

Phase 3 — Block Storage (Months 5-7)

Deliverables:
  ✓ VolumeService in MDS
  ✓ Volume extent map with thin provisioning
  ✓ COW snapshots
  ✓ WAL for write ordering on DSS
  ✓ crates/iscsi-target/ — iSCSI target daemon

Milestone: boot a VM from RustFS block volume

Acceptance test:
  $ rustfs-admin volume create --name=vm-boot --size=50G --thin
  $ rustfs-iscsi --mds=mds1:9100 --volume=vm-boot --lun=0 &
  $ qemu-system-x86_64 \
      -drive file=iscsi://localhost/iqn.2026-06.com.rustfs:vm-boot/0 \
      -m 4G
  # VM boots from RustFS block volume
  # Verify: snapshot, clone, resize operations

Phase 4 — Production (Months 7-10)

Deliverables:
  ✓ HA MDS with Raft (3-node or 5-node cluster)
  ✓ MDS sharding (subtree partitioning)
  ✓ ecstore decomposition (monolith → sub-crates)
  ✓ Kubernetes CSI driver (RWX for POSIX, RWO for block)
  ✓ NVMe-oF target for high-performance block

Milestone: pass production readiness review

Phase 5 — Performance (Months 10-14)

Deliverables:
  ✓ RDMA transport for chunk reads (DSS ↔ client)
  ✓ GPUDirect Storage integration (bypass CPU on reads)
  ✓ io_uring FUSE passthrough (bypass FUSE kernel overhead)
  ✓ Client write-back caching with lease coherency
  ✓ Automated tiering (hot replicated → cold EC)

Milestone: match Lustre/BeeGFS on IOR and mdtest benchmarks

11. Compatibility Matrix

Feature                  S3 API          POSIX          Block
───────────────────────  ──────────────  ─────────────  ─────────────
Create file/object       PUT Object      creat()/open() create-vol
Read                     GET Object      read()         SCSI READ
Write                    PUT Object      write()        SCSI WRITE
Delete                   DELETE Object   unlink()       delete-vol
List                     ListObjects     readdir()      list-vols
Rename                   COPY+DELETE     rename()       N/A
Permissions              Bucket Policy   chmod/chown    LUN masking
                         + IAM           + POSIX ACL
Versioning               S3 Versioning   (via snapshots) COW snapshots
Locking                  S3 Object Lock  flock/fcntl    iSCSI reservations
Notifications            S3 Events       inotify (future)  N/A
Encryption               SSE-S3/SSE-KMS  (inherited)    (inherited)
Compression              (inherited)     (inherited)    (disabled)
Quotas                   Bucket quota    Dir quota      Vol capacity
Replication              Bucket repl.    (inherited)    Vol repl factor

12. Open Design Questions

ID   Question                                    Decision needed by
──── ──────────────────────────────────────────── ─────────────────
DQ1  Should MDS store chunk maps inline with      Phase 1 start
     inodes (simpler) or in a separate CF
     (better for large files with millions of
     chunks)?

DQ2  For the S3 ↔ POSIX bridge, should we        Phase 2 start
     support hard links between S3 "objects"?
     (S3 has no link concept)

DQ3  For block volumes, should we use a separate  Phase 3 start
     WAL per volume or a shared WAL per DSS node?

DQ4  RDMA transport: use existing tonic/gRPC      Phase 5 start
     with RDMA transport, or implement custom
     zero-copy RDMA verbs?

DQ5  Should the FUSE client support direct        Phase 2 start
     kernel bypass (io_uring passthrough or
     CUSE) from day 1, or add it later?

DQ6  MDS sharding strategy: static subtree        Phase 4 start
     assignment or dynamic migration (like
     CephFS)?

Appendix A: Reference Architectures

System       MDS                  Data Path           Block
───────────  ───────────────────  ──────────────────  ────────────
Lustre       Separate MDS (MDD    OST servers, file   N/A
             on ldiskfs/ZFS)      striping, LDLM

BeeGFS       Separate MDS         OST servers,        N/A (BeeGFS
             (BuddyMirror HA)     file striping       Mon only)

CephFS       MDS cluster          RADOS objects        Ceph RBD
             (Raft-like)          (EC or replicated)  (same RADOS)

JuiceFS      External DB          Any S3-compatible   N/A
             (Redis/TiKV/MySQL)   object store

RustFS       Embedded in          Erasure-coded       N/A
(current)    ecstore (no          shards on local
             separate MDS)        disks

RustFS       Separate MDS         Chunk service on    Volume manager
(proposed)   (RocksDB + Raft)     existing DSS nodes  + iSCSI/NVMe-oF
                                  (EC or replicated)  on same DSS

End of specification. This document should be kept in sync with implementation as the codebase evolves. Update the version number and revision date on every significant change.