RustFS Unified Storage Platform
RustFS Unified Storage Platform
Feature Specification & High-Level Design
Document: RUSTFS-FEAT-001
Status: DRAFT
Version: 0.1.0
Authors: Architecture Team
Created: 2026-06-05
Audience: Engineering (implementation-ready)
Codebase: github.com/rustfs/rustfs
0. Document Purpose
This specification defines the design of RustFS Unified Storage — a single storage platform that exposes file (POSIX), object (S3), and block access to the same data through one namespace. It is written for engineers who will implement the system. Every section maps to concrete crates, types, protobuf definitions, and code paths in the RustFS codebase.
The target workloads are:
- AI training / inference: POSIX
open → fseek → freadfrom hundreds of GPU nodes reading the same dataset concurrently, with high aggregate bandwidth (100+ GB/s cluster) and low per-op latency (<1 ms metadata, <5 ms first-byte data). - VM / bare-metal boot disks: 4K random read/write block I/O over iSCSI or NVMe-oF, with sub-millisecond latency, thin provisioning, and COW snapshots.
- Object storage: existing S3-compatible API, unchanged, fully interoperable with the file and block namespaces.
1. Unified Namespace Model
1.1 Design Principle: One Namespace, Three Projections
Every piece of data in the cluster lives under a single hierarchical namespace
rooted at /. The three access protocols project different views of this
namespace:
┌──────────────────────┐
│ Unified Namespace │
│ / │
│ ┌─────┴──────┐ │
│ │ volumes/ │ ... │
│ │ data/ │ │
│ │ buckets/ │ │
│ └────────────┘ │
└──────┬───┬───┬────────┘
│ │ │
┌───────────────┘ │ └───────────────┐
▼ ▼ ▼
┌─────────────┐ ┌──────────────┐ ┌──────────────┐
│ POSIX View │ │ S3 View │ │ Block View │
│ │ │ │ │ │
│ mount -t │ │ GET /bucket │ │ /dev/rustfs/ │
│ rustfs │ │ /key HTTP/1.1│ │ vol-001 │
│ /mnt/rustfs │ │ │ │ │
│ │ │ bucket = dir │ │ volume = │
│ ls, cat, │ │ key = path │ │ extent- │
│ fseek, fread │ │ under bucket │ │ mapped file │
└─────────────┘ └──────────────┘ └──────────────┘
1.2 Namespace Hierarchy
/ # root inode (ino=1)
├── buckets/ # S3 bucket mount root
│ ├── training-data/ # S3 bucket → directory
│ │ ├── imagenet/
│ │ │ ├── shard-0001.tar # regular file (S3 object)
│ │ │ └── shard-0002.tar
│ │ └── llm/
│ │ └── tokens.bin
│ └── model-checkpoints/
│ └── epoch-042/
│ └── weights.safetensors
├── volumes/ # block volume mount root
│ ├── vm-disk-001 # block device special file
│ ├── vm-disk-002
│ └── db-wal-001
└── scratch/ # general POSIX workspace
└── user-alice/
└── experiment-7/
1.3 Identity Mapping Rules
Every entity in the namespace is an inode. The access protocol determines how the inode is presented:
Inode kind POSIX view S3 view Block view
─────────────────── ─────────────────────── ──────────────────────── ──────────────
RegularFile regular file object (bucket = parent N/A
(open/read/write/seek) dir under /buckets/)
Directory directory bucket (if depth=1 N/A
(readdir/mkdir/rmdir) under /buckets/) or
common prefix
Symlink symlink not exposed N/A
BlockDevice /dev/rustfs/vol-xxx not exposed LUN / NVMe
(can be opened with namespace
O_RDWR for raw I/O)
1.4 Path ↔ S3 Key Translation
POSIX path: /buckets/training-data/imagenet/shard-0001.tar
S3 address: bucket = "training-data", key = "imagenet/shard-0001.tar"
Rule:
bucket_name = first path component under /buckets/
object_key = remaining path components joined by "/"
Reverse (S3 → POSIX):
S3 PUT training-data/new-folder/file.bin
→ MDS: mkdir_p(/buckets/training-data/new-folder/)
→ MDS: create(/buckets/training-data/new-folder/file.bin)
2. System Architecture
2.1 Component Topology
┌─────────────────────────────────────────────────────────────────────┐
│ CLIENT TIER │
│ │
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────────────┐ │
│ │ FUSE │ │ S3 SDK │ │ iSCSI │ │ NVMe-oF │ │
│ │ Client │ │ Client │ │ Initiator│ │ Host Driver │ │
│ │(rustfs- │ │(any S3 │ │(Linux │ │(kernel nvme-tcp │ │
│ │ mount) │ │ client) │ │ open- │ │ or nvme-rdma) │ │
│ │ │ │ │ │ iscsi) │ │ │ │
│ └────┬─────┘ └────┬─────┘ └────┬─────┘ └────────┬─────────┘ │
│ │ │ │ │ │
│ │ gRPC │ HTTP │ iSCSI │ NVMe-oF │
└───────┼──────────────┼──────────────┼─────────────────┼─────────────┘
│ │ │ │
═══════════════════════════════════════════════════════════════════════
│ │ │ │
┌───────┼──────────────┼──────────────┼─────────────────┼─────────────┐
│ │ GATEWAY TIER │ │ │
│ │ │ │ │ │
│ │ ┌────┴─────┐ ┌────┴────┐ ┌──────┴──────┐ │
│ │ │ S3 API │ │ iSCSI │ │ NVMe-oF │ │
│ │ │ Server │ │ Target │ │ Target │ │
│ │ │ (existing│ │(rustfs- │ │(rustfs- │ │
│ │ │ rustfs │ │ iscsi) │ │ nvmeof) │ │
│ │ │ binary) │ │ │ │ │ │
│ │ └────┬─────┘ └────┬────┘ └──────┬──────┘ │
│ │ │ │ │ │
└───────┼──────────────┼──────────────┼─────────────────┼─────────────┘
│ │ │ │
▼ ▼ ▼ ▼
┌─────────────────────────────────────────────────────────────────────┐
│ METADATA TIER (MDS) │
│ │
│ ┌────────────────────────────────────────────────────────────────┐ │
│ │ MDS Cluster (Raft) │ │
│ │ │ │
│ │ ┌──────────┐ ┌──────────┐ ┌──────────┐ │ │
│ │ │ MDS-0 │──│ MDS-1 │──│ MDS-2 │ (3 or 5 nodes) │ │
│ │ │ (leader) │ │(follower)│ │(follower)│ │ │
│ │ └──────────┘ └──────────┘ └──────────┘ │ │
│ │ │ │ │
│ │ ┌────┴────────────────────────────────────────────┐ │ │
│ │ │ RocksDB: inodes, dentries, chunk maps, │ │ │
│ │ │ volumes, leases, locks │ │ │
│ │ └─────────────────────────────────────────────────┘ │ │
│ └────────────────────────────────────────────────────────────────┘ │
│ │
└─────────────────────────────────────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────────────┐
│ DATA TIER (DSS — Data Storage Servers) │
│ │
│ ┌─────────┐ ┌─────────┐ ┌─────────┐ ┌─────────┐ │
│ │ DSS-0 │ │ DSS-1 │ │ DSS-2 │ ... │ DSS-N │ │
│ │ │ │ │ │ │ │ │ │
│ │ chunk │ │ chunk │ │ chunk │ │ chunk │ │
│ │ service │ │ service │ │ service │ │ service │ │
│ │ (gRPC) │ │ (gRPC) │ │ (gRPC) │ │ (gRPC) │ │
│ │ │ │ │ │ │ │ │ │
│ │ ┌─────┐ │ │ ┌─────┐ │ │ ┌─────┐ │ │ ┌─────┐ │ │
│ │ │disk1│ │ │ │disk1│ │ │ │disk1│ │ │ │disk1│ │ │
│ │ │disk2│ │ │ │disk2│ │ │ │disk2│ │ │ │disk2│ │ │
│ │ │ ... │ │ │ │ ... │ │ │ │ ... │ │ │ │ ... │ │ │
│ │ │diskM│ │ │ │diskM│ │ │ │diskM│ │ │ │diskM│ │ │
│ │ └─────┘ │ │ └─────┘ │ │ └─────┘ │ │ └─────┘ │ │
│ └─────────┘ └─────────┘ └─────────┘ └─────────┘ │
│ │
└─────────────────────────────────────────────────────────────────────┘
2.2 Deployment Modes
Mode MDS DSS Use case
──────────────── ─────────────── ─────────────── ─────────────────
Standalone embedded in same process Dev / edge / IoT
rustfs binary (single node)
Small cluster 1 MDS (no HA) 3-8 DSS nodes Small AI lab
(< 50 GPUs)
Production 3-5 MDS Raft 8-128 DSS nodes AI training center
cluster (50-1000 GPUs)
Hyperscale Sharded MDS 128+ DSS nodes Cloud-scale
(DNE: subtree (1000+ GPUs,
partitioning) PB-scale)
2.3 Data Flow Summary
┌──────────────────────────────────┐
│ CLIENT REQUEST │
└────────────────┬─────────────────┘
│
┌────────────────┴─────────────────┐
│ WHAT KIND OF OP? │
└──┬──────────┬──────────┬─────────┘
│ │ │
metadata op data read data write
│ │ │
▼ │ │
┌──────────┐ │ │
│ MDS │ │ │
│ (lookup, │ │ │
│ stat, │◄────┤ │
│ mkdir, │ │ │
│ create) │ │ │
└────┬─────┘ │ │
│ │ │
returns chunk │ │
map + lease │ │
│ │ │
▼ ▼ ▼
┌────────────────────────────────┐
│ DSS (Data Storage Servers) │
│ │
│ read: parallel chunk fetch │
│ write: parallel chunk write │
│ + MDS commit │
└─────────────────────────────────┘
3. Core Data Structures
3.1 Inode
The universal identity for every object in the unified namespace.
// crate: rustfs-inode
// file: crates/inode/src/types.rs
/// 64-bit inode identifier.
///
/// Layout:
/// bits [63..48] = MDS partition ID (supports 65536 MDS shards)
/// bits [47..0] = local sequence number (281 trillion inodes per shard)
///
/// Partition 0 is the default single-MDS partition.
/// ROOT_INODE = 0x0000_0000_0000_0001
pub type InodeId = u64;
pub const ROOT_INODE: InodeId = 1;
/// Inode kinds — determines which protocol views can access this inode.
#[derive(Clone, Debug, PartialEq, Eq, Serialize, Deserialize)]
pub enum InodeKind {
RegularFile,
Directory,
Symlink { target: String },
BlockVolume { volume_id: VolumeId },
}
/// Full POSIX inode attributes.
///
/// Stored in MDS RocksDB under CF "inodes", key = big-endian InodeId bytes.
/// Serialized as MessagePack for compact storage and fast decode.
#[derive(Clone, Debug, Serialize, Deserialize)]
pub struct InodeAttr {
pub ino: InodeId,
pub kind: InodeKind,
pub size: u64,
pub blocks: u64, // in 512-byte units (for stat.st_blocks)
pub atime: SystemTime,
pub mtime: SystemTime,
pub ctime: SystemTime,
pub birthtime: SystemTime,
pub perm: u16, // rwxrwxrwx (9 bits)
pub uid: u32,
pub gid: u32,
pub nlink: u32,
pub rdev: u32,
pub flags: u32,
pub generation: u64, // for NFS export and FUSE
// Storage policy (set at creation, changeable via admin API)
pub storage_policy: StoragePolicy,
// S3 compatibility fields (populated when accessed via S3)
pub content_type: Option<String>,
pub content_encoding: Option<String>,
pub etag: Option<String>,
pub s3_metadata: BTreeMap<String, String>, // x-amz-meta-*
pub s3_tags: BTreeMap<String, String>, // S3 object tags
pub s3_storage_class: Option<String>,
}
/// Storage policy determines how data chunks are protected.
///
/// Selected per-file or per-directory (inherited by children).
/// AI training data → EC striping (write-once, read-many, max throughput)
/// VM block volumes → replication (random 4K writes, low latency)
/// General POSIX → configurable (default EC for large, replica for small)
#[derive(Clone, Debug, Serialize, Deserialize)]
pub enum StoragePolicy {
/// Reed-Solomon erasure coding across N data + M parity shards.
/// Best for: large files, sequential I/O, write-once-read-many.
/// Inherited from existing ecstore EC engine.
ErasureCoded {
data_shards: u32, // e.g., 4
parity_shards: u32, // e.g., 2
chunk_size: u64, // e.g., 4 MiB
},
/// N-way replication.
/// Best for: small files, random I/O, block volumes.
Replicated {
factor: u32, // e.g., 3
chunk_size: u64, // e.g., 4 MiB (file) or 4 MiB (block)
},
/// Tiered: hot data replicated, cold data EC'd.
/// Best for: mixed workloads with lifecycle management.
Tiered {
hot: Box<StoragePolicy>,
cold: Box<StoragePolicy>,
hot_threshold_days: u32,
},
}
impl Default for StoragePolicy {
fn default() -> Self {
StoragePolicy::ErasureCoded {
data_shards: 4,
parity_shards: 2,
chunk_size: 4 * 1024 * 1024, // 4 MiB
}
}
}
3.2 Directory Entry
// crate: rustfs-inode
// file: crates/inode/src/dentry.rs
/// A single directory entry.
///
/// Stored in MDS RocksDB under CF "dentries".
/// Key = big-endian parent InodeId ++ name bytes (allows prefix scan)
/// Value = DirEntry as MessagePack
#[derive(Clone, Debug, Serialize, Deserialize)]
pub struct DirEntry {
pub name: String,
pub ino: InodeId,
pub kind: InodeKind, // cached from inode (avoids extra lookup)
}
/// Batch directory listing result for readdir.
#[derive(Clone, Debug, Serialize, Deserialize)]
pub struct DirListing {
pub parent: InodeId,
pub entries: Vec<DirEntry>,
pub has_more: bool,
pub next_cursor: Option<String>, // for paginated readdir
}
3.3 Chunk Map
// crate: rustfs-inode
// file: crates/inode/src/chunkmap.rs
/// Globally unique chunk identifier.
///
/// Format: "{dss_node_id}-{disk_id}-{local_sequence}"
/// Example: "dss03-disk07-000000004a2f"
pub type ChunkId = String;
/// Which DSS node owns a chunk.
pub type NodeId = u32;
/// Physical location of one chunk replica or EC shard.
#[derive(Clone, Debug, Serialize, Deserialize)]
pub struct ChunkLocation {
pub chunk_id: ChunkId,
pub node_id: NodeId,
pub disk_id: u32,
pub version: u64, // monotonic, for COW detection
pub status: ChunkStatus,
}
#[derive(Clone, Debug, Serialize, Deserialize)]
pub enum ChunkStatus {
Active,
Healing,
Stale,
}
/// Maps a file's byte ranges to chunk storage locations.
///
/// Stored in MDS RocksDB under CF "chunkmaps".
/// Key = big-endian InodeId
/// Value = ChunkMap as MessagePack
///
/// Design: BTreeMap indexed by chunk_index (= file_offset / chunk_size).
/// This gives O(log N) range queries for any [offset, offset+len) read/write.
///
/// For EC policy: each chunk_index maps to N+M ChunkLocations
/// (one per EC shard across different DSS nodes).
/// For replica policy: each chunk_index maps to R ChunkLocations
/// (one per replica across different DSS nodes).
#[derive(Clone, Debug, Serialize, Deserialize)]
pub struct ChunkMap {
pub ino: InodeId,
pub chunk_size: u64,
pub policy: StoragePolicy,
pub chunks: BTreeMap<u64, Vec<ChunkLocation>>,
}
impl ChunkMap {
/// Resolve which chunks cover byte range [offset, offset+length).
///
/// Returns: Vec<(chunk_index, offset_within_chunk, read_length, locations)>
///
/// This is the HOT PATH for every POSIX read and block read.
/// Must be O(log N + K) where K = number of chunks in the range.
pub fn resolve_range(
&self,
offset: u64,
length: u64,
) -> Vec<ChunkRangeEntry> {
let start_idx = offset / self.chunk_size;
let end_idx = (offset + length).saturating_sub(1) / self.chunk_size;
self.chunks.range(start_idx..=end_idx).map(|(idx, locs)| {
let chunk_start = idx * self.chunk_size;
let local_offset = if offset > chunk_start {
offset - chunk_start
} else { 0 };
let local_end = ((idx + 1) * self.chunk_size).min(offset + length);
let local_length = local_end - chunk_start - local_offset;
ChunkRangeEntry {
chunk_index: *idx,
offset_in_chunk: local_offset,
length: local_length,
locations: locs.clone(),
}
}).collect()
}
/// Total number of chunks (for capacity accounting).
pub fn chunk_count(&self) -> u64 {
self.chunks.len() as u64
}
}
#[derive(Clone, Debug)]
pub struct ChunkRangeEntry {
pub chunk_index: u64,
pub offset_in_chunk: u64,
pub length: u64,
pub locations: Vec<ChunkLocation>,
}
3.4 Block Volume
// crate: rustfs-inode
// file: crates/inode/src/volume.rs
pub type VolumeId = String; // e.g., "vol-00af39b2"
pub type SnapshotId = String;
/// Block volume descriptor.
///
/// A volume is represented in the unified namespace as:
/// InodeKind::BlockVolume { volume_id }
/// at path /volumes/{volume_name}
///
/// The volume's data is stored as an extent-mapped set of chunks,
/// reusing the same ChunkMap structure but with different access semantics:
/// - Block I/O is always aligned to volume.block_size
/// - Writes are partial (4K granularity within 4 MiB chunks)
/// - Flush / write-barrier semantics are enforced
///
/// Stored in MDS RocksDB under CF "volumes", key = VolumeId.
#[derive(Clone, Debug, Serialize, Deserialize)]
pub struct Volume {
pub id: VolumeId,
pub name: String, // human-readable
pub ino: InodeId, // inode in /volumes/
pub capacity: u64, // logical size in bytes
pub block_size: u32, // 512 or 4096
pub chunk_size: u64, // 4 MiB default
pub provisioning: VolumeProvisioning,
pub state: VolumeState,
pub snapshots: Vec<SnapshotId>,
pub created_at: SystemTime,
pub chunk_map: ChunkMap, // reuses file chunk map!
// Storage: always replicated (not EC) for random write performance
pub replica_factor: u32, // default 3
}
#[derive(Clone, Debug, Serialize, Deserialize)]
pub enum VolumeProvisioning {
Thin, // allocate on first write (default)
Thick, // pre-allocate all chunks at creation
}
#[derive(Clone, Debug, Serialize, Deserialize)]
pub enum VolumeState {
Available,
Attached { target: String }, // which host / initiator
Snapshotting,
Deleting,
}
/// COW snapshot of a volume.
///
/// Shares the parent's chunk map at snapshot time.
/// Subsequent writes to parent or snapshot COW the affected chunks.
#[derive(Clone, Debug, Serialize, Deserialize)]
pub struct VolumeSnapshot {
pub id: SnapshotId,
pub volume_id: VolumeId,
pub created_at: SystemTime,
pub capacity: u64,
pub chunk_map: ChunkMap, // frozen copy at snapshot time
pub cow_version: u64, // version counter for COW tracking
}
3.5 Leases and Locks
// crate: rustfs-inode
// file: crates/inode/src/lease.rs
pub type ClientId = u64;
pub type SessionId = u64;
/// Client lease — enables client-side caching with coherency.
///
/// Inspired by Lustre LDLM and NFSv4 delegations.
///
/// The MDS grants leases; clients cache data/metadata locally.
/// When a conflicting operation arrives, MDS revokes the lease
/// (callback to client) before proceeding.
///
/// This is the KEY mechanism for AI training performance:
/// hundreds of readers get ReadData leases on training files →
/// each reader caches chunks locally → no MDS/DSS traffic on
/// repeated epoch reads → GPU feed rate maximized.
#[derive(Clone, Debug, Serialize, Deserialize)]
pub struct Lease {
pub id: u64,
pub ino: InodeId,
pub client_id: ClientId,
pub session_id: SessionId,
pub kind: LeaseKind,
pub granted_at: SystemTime,
pub expires_at: SystemTime, // must heartbeat to renew
}
#[derive(Clone, Debug, PartialEq, Eq, Serialize, Deserialize)]
pub enum LeaseKind {
/// Client can cache inode attributes (stat cache).
/// Revoked when any client modifies attributes.
ReadMeta,
/// Client can cache file data in its page cache.
/// Revoked when any client writes to this file.
/// Multiple ReadData leases can coexist.
ReadData,
/// Client has exclusive write access and can cache writes.
/// Only one WriteData lease per inode.
/// Revoked when another client requests any access.
WriteData,
/// Client knows the chunk map layout and can issue direct
/// chunk reads to DSS without re-checking MDS.
/// Revoked when file is re-striped, truncated, or extended.
Layout,
}
/// POSIX byte-range lock (flock / fcntl).
///
/// Stored in MDS memory (not persisted — locks die with the MDS).
/// For HA MDS, locks are part of the Raft state machine.
#[derive(Clone, Debug, Serialize, Deserialize)]
pub struct ByteRangeLock {
pub ino: InodeId,
pub owner: (ClientId, u64), // (client, pid or lock_owner)
pub kind: RangeLockKind,
pub start: u64,
pub end: u64, // u64::MAX = to EOF
}
#[derive(Clone, Debug, PartialEq, Eq, Serialize, Deserialize)]
pub enum RangeLockKind {
Read, // F_RDLCK — shared
Write, // F_WRLCK — exclusive
}
3.6 MDS RocksDB Schema
Column Family Key Value (msgpack)
───────────────────── ───────────────────────────── ───────────────────────
"inodes" BE(inode_id) InodeAttr
"dentries" BE(parent_ino) ++ name_bytes DirEntry
"chunkmaps" BE(inode_id) ChunkMap
"volumes" volume_id_bytes Volume
"snapshots" snapshot_id_bytes VolumeSnapshot
"xattrs" BE(inode_id) ++ xattr_name xattr_value bytes
"symlinks" BE(inode_id) target_path string
"s3multipart" upload_id_bytes MultipartUploadState
Meta keys (default CF):
"next_inode" → u64 (atomic inode counter)
"cluster_id" → UUID
"mds_version" → u32
(BE = big-endian encoding for correct byte-order sorting in RocksDB)
4. MDS Service Definition
4.1 gRPC API
// file: proto/mds.proto
syntax = "proto3";
package rustfs.mds;
// ═══════════════════════════════════════════════════════
// SESSION MANAGEMENT
// ═══════════════════════════════════════════════════════
service SessionService {
// Client registration. Returns client_id and session_id.
// Client must call Heartbeat periodically (default: 30s).
// If heartbeat lapses, MDS revokes all leases and locks.
rpc OpenSession(OpenSessionReq) returns (OpenSessionResp);
rpc Heartbeat(HeartbeatReq) returns (HeartbeatResp);
rpc CloseSession(CloseSessionReq) returns (CloseSessionResp);
}
message OpenSessionReq {
string client_hostname = 1;
string client_version = 2;
uint32 max_read_ahead_mb = 3; // client's read-ahead capability
}
message OpenSessionResp {
uint64 client_id = 1;
uint64 session_id = 2;
uint32 heartbeat_interval_ms = 3; // server-chosen interval
repeated string dss_endpoints = 4; // DSS node addresses for data path
}
// ═══════════════════════════════════════════════════════
// NAMESPACE OPERATIONS
// ═══════════════════════════════════════════════════════
service NamespaceService {
// Path resolution
rpc Lookup(LookupReq) returns (LookupResp);
// Inode attributes
rpc GetAttr(GetAttrReq) returns (GetAttrResp);
rpc SetAttr(SetAttrReq) returns (SetAttrResp);
// File lifecycle
rpc Create(CreateReq) returns (CreateResp);
rpc Unlink(UnlinkReq) returns (UnlinkResp);
rpc Rename(RenameReq) returns (RenameResp);
rpc Link(LinkReq) returns (LinkResp);
rpc Symlink(SymlinkReq) returns (SymlinkResp);
// Directory operations
rpc MkDir(MkDirReq) returns (MkDirResp);
rpc RmDir(RmDirReq) returns (RmDirResp);
rpc ReadDir(ReadDirReq) returns (ReadDirResp);
// Extended attributes
rpc GetXAttr(GetXAttrReq) returns (GetXAttrResp);
rpc SetXAttr(SetXAttrReq) returns (SetXAttrResp);
rpc ListXAttr(ListXAttrReq) returns (ListXAttrResp);
rpc RemoveXAttr(RemoveXAttrReq) returns (RemoveXAttrResp);
}
message LookupReq {
uint64 session_id = 1;
uint64 parent_ino = 2;
string name = 3;
}
message LookupResp {
uint64 ino = 1;
InodeAttr attr = 2;
uint64 generation = 3;
}
message CreateReq {
uint64 session_id = 1;
uint64 parent_ino = 2;
string name = 3;
uint32 mode = 4; // permission bits
uint32 uid = 5;
uint32 gid = 6;
uint32 flags = 7; // O_EXCL, O_TRUNC, etc.
StoragePolicy policy = 8; // optional override
}
message CreateResp {
uint64 ino = 1;
InodeAttr attr = 2;
uint64 generation = 3;
Lease lease = 4; // auto-granted write lease
}
message ReadDirReq {
uint64 session_id = 1;
uint64 ino = 2;
uint64 offset = 3; // cookie for pagination
uint32 limit = 4; // max entries to return
bool plus = 5; // readdirplus: include attrs
}
message ReadDirResp {
repeated DirEntryProto entries = 1;
bool has_more = 2;
uint64 next_offset = 3;
}
// ═══════════════════════════════════════════════════════
// DATA LAYOUT (CHUNK MAP) OPERATIONS
// ═══════════════════════════════════════════════════════
service DataLayoutService {
// Get chunk map for a file (used before data I/O)
rpc GetChunkMap(GetChunkMapReq) returns (GetChunkMapResp);
// Allocate new chunks for writes (MDS picks DSS nodes)
rpc AllocChunks(AllocChunksReq) returns (AllocChunksResp);
// Commit chunk writes (update chunk map after successful DSS writes)
rpc CommitWrite(CommitWriteReq) returns (CommitWriteResp);
// Truncate file to new size (free chunks beyond new size)
rpc Truncate(TruncateReq) returns (TruncateResp);
}
message GetChunkMapReq {
uint64 session_id = 1;
uint64 ino = 2;
uint64 offset = 3; // start of range (0 = entire file)
uint64 length = 4; // 0 = to EOF
}
message GetChunkMapResp {
ChunkMapProto chunk_map = 1;
Lease layout_lease = 2; // optional layout lease
}
message AllocChunksReq {
uint64 session_id = 1;
uint64 ino = 2;
uint64 start_index = 3; // first chunk index to allocate
uint32 count = 4; // number of chunks to allocate
StoragePolicy policy = 5;
}
message AllocChunksResp {
repeated ChunkAllocation allocations = 1;
}
message ChunkAllocation {
uint64 chunk_index = 1;
repeated ChunkLocationProto locations = 2; // where to write
}
message CommitWriteReq {
uint64 session_id = 1;
uint64 ino = 2;
uint64 new_size = 3; // file size after this write
int64 mtime_ns = 4; // new mtime
repeated ChunkCommit chunks = 5;
}
message ChunkCommit {
uint64 chunk_index = 1;
repeated ChunkLocationProto locations = 2;
uint64 bytes_written = 3;
}
// ═══════════════════════════════════════════════════════
// LEASE AND LOCK OPERATIONS
// ═══════════════════════════════════════════════════════
service LeaseService {
rpc AcquireLease(AcquireLeaseReq) returns (AcquireLeaseResp);
rpc RenewLease(RenewLeaseReq) returns (RenewLeaseResp);
rpc ReleaseLease(ReleaseLeaseReq) returns (ReleaseLeaseResp);
// Server → Client callback for lease revocation.
// Client implements this as a gRPC server endpoint.
// MDS calls this when a conflicting operation requires revocation.
rpc RevokeLease(RevokeLeaseReq) returns (RevokeLeaseResp);
// POSIX byte-range locks
rpc LockRange(LockRangeReq) returns (LockRangeResp);
rpc UnlockRange(UnlockRangeReq) returns (UnlockRangeResp);
rpc TestLock(TestLockReq) returns (TestLockResp);
}
// ═══════════════════════════════════════════════════════
// BLOCK VOLUME OPERATIONS
// ═══════════════════════════════════════════════════════
service VolumeService {
rpc CreateVolume(CreateVolumeReq) returns (CreateVolumeResp);
rpc DeleteVolume(DeleteVolumeReq) returns (DeleteVolumeResp);
rpc ResizeVolume(ResizeVolumeReq) returns (ResizeVolumeResp);
rpc AttachVolume(AttachVolumeReq) returns (AttachVolumeResp);
rpc DetachVolume(DetachVolumeReq) returns (DetachVolumeResp);
rpc SnapshotVolume(SnapshotReq) returns (SnapshotResp);
rpc CloneVolume(CloneVolumeReq) returns (CloneVolumeResp);
rpc ListVolumes(ListVolumesReq) returns (ListVolumesResp);
rpc GetVolumeChunkMap(GetVolMapReq) returns (GetVolMapResp);
}
message CreateVolumeReq {
string name = 1;
uint64 capacity_bytes = 2;
uint32 block_size = 3; // 512 or 4096
string provisioning = 4; // "thin" or "thick"
uint32 replica_factor = 5; // default 3
}
4.2 S3 API Adaptation
The existing rustfs S3 API server calls into MDS instead of directly into
ecstore for namespace operations:
// file: rustfs/src/storage/unified.rs (NEW — replaces parts of ecfs.rs)
/// Unified storage adapter.
///
/// Translates S3 API calls into MDS + DSS operations.
/// Installed as the ObjectLayer implementation when unified mode is enabled.
pub struct UnifiedObjectLayer {
mds: MdsClient,
dss_pool: DssConnectionPool,
s3_compat: S3CompatLayer, // handles S3-specific quirks
}
impl ObjectLayer for UnifiedObjectLayer {
async fn put_object(&self, bucket: &str, key: &str, data: Stream, opts: PutOpts)
-> Result<ObjectInfo>
{
// 1. Resolve path: /buckets/{bucket}/{key}
let parent_ino = self.resolve_s3_path(bucket, key).await?;
// 2. Create inode via MDS (or overwrite existing)
let create_resp = self.mds.create(CreateReq {
parent_ino,
name: key_last_component(key),
mode: 0o644,
policy: self.policy_for_bucket(bucket),
..Default::default()
}).await?;
// 3. Allocate chunks via MDS
let chunk_count = (data.content_length + chunk_size - 1) / chunk_size;
let alloc = self.mds.alloc_chunks(AllocChunksReq {
ino: create_resp.ino,
count: chunk_count as u32,
..Default::default()
}).await?;
// 4. Write data to DSS nodes (parallel)
let commits = self.write_chunks_parallel(&alloc, data).await?;
// 5. Commit to MDS
self.mds.commit_write(CommitWriteReq {
ino: create_resp.ino,
new_size: data.content_length,
chunks: commits,
..Default::default()
}).await?;
// 6. Return S3-compatible ObjectInfo
Ok(self.s3_compat.to_object_info(&create_resp.attr))
}
async fn get_object(&self, bucket: &str, key: &str, range: Option<HttpRange>)
-> Result<GetObjectResp>
{
// 1. Resolve path → inode
let ino = self.resolve_s3_key(bucket, key).await?;
// 2. Get chunk map from MDS
let (offset, length) = range.unwrap_or((0, u64::MAX));
let cm = self.mds.get_chunk_map(ino, offset, length).await?;
// 3. Read chunks from DSS (parallel — reuses existing ecstore logic)
let stream = self.read_chunks_parallel(&cm, offset, length).await?;
Ok(GetObjectResp { stream, attr: cm.attr })
}
async fn list_objects(&self, bucket: &str, prefix: &str, opts: ListOpts)
-> Result<ListObjectsResp>
{
// Translate to MDS readdir on the resolved directory.
// If prefix contains "/", walk directory tree.
// Much faster than scanning all keys!
let dir_ino = self.resolve_s3_prefix(bucket, prefix).await?;
let listing = self.mds.read_dir(dir_ino, opts.marker, opts.max_keys).await?;
Ok(self.s3_compat.to_list_response(listing, prefix, opts.delimiter))
}
}
5. DSS Chunk Service
5.1 gRPC API
// file: proto/chunk.proto
syntax = "proto3";
package rustfs.chunk;
service ChunkService {
// Write data to a chunk at an offset.
//
// For replicated chunks: direct byte-range write.
// For EC chunks: if write covers full EC block(s), encode + write shards.
// if partial, read-modify-write the affected block.
rpc WriteChunk(stream WriteChunkReq) returns (WriteChunkResp);
// Read data from a chunk at an offset.
//
// This reuses and extends the existing rustfs RPC:
// /rustfs/rpc/read_file_stream?disk=...&offset=...&length=...
rpc ReadChunk(ReadChunkReq) returns (stream ReadChunkResp);
// Batch read: read multiple chunk ranges in one RPC.
// Reduces round trips for striped reads.
rpc ReadChunkBatch(ReadChunkBatchReq) returns (stream ReadChunkBatchResp);
// Ensure chunk data is durable (fsync).
rpc FlushChunk(FlushChunkReq) returns (FlushChunkResp);
// Truncate or delete chunk.
rpc TruncateChunk(TruncateChunkReq) returns (TruncateChunkResp);
rpc DeleteChunks(DeleteChunksReq) returns (DeleteChunksResp);
// Status
rpc StatChunk(StatChunkReq) returns (StatChunkResp);
}
message WriteChunkReq {
string chunk_id = 1;
uint64 offset = 2; // byte offset within chunk
bytes data = 3; // payload (streamed for large writes)
bool sync = 4; // fsync after write
bool create = 5; // create chunk if not exists
}
message WriteChunkResp {
uint64 bytes_written = 1;
string checksum = 2; // new checksum after write
}
message ReadChunkReq {
string chunk_id = 1;
uint64 offset = 2;
uint64 length = 3;
bool direct_io = 4; // bypass OS page cache
bool verify = 5; // verify bitrot checksum
}
message ReadChunkResp {
bytes data = 1; // streamed
uint64 offset = 2; // offset of this chunk within stream
}
message ReadChunkBatchReq {
repeated ReadChunkReq reads = 1;
}
message ReadChunkBatchResp {
uint32 index = 1; // which request in the batch
bytes data = 2; // streamed
}
5.2 DSS Chunk Storage on Disk
Physical disk layout per DSS node:
/{volume_mount}/
├── .rustfs.sys/ # cluster metadata (existing)
├── .chunks/ # chunk data (NEW)
│ ├── 00/ # first 2 hex chars of chunk_id
│ │ ├── 00a3f9b2.dat # chunk data file
│ │ ├── 00a3f9b2.meta # chunk metadata (checksum, size)
│ │ └── 00b1cc07.dat
│ ├── 01/
│ │ └── ...
│ └── ff/
│ └── ...
├── {bucket}/ # S3 legacy object data (existing)
│ └── {key_hash}/
│ ├── xl.meta
│ └── part.1
└── .journal/ # write-ahead log for block volumes
├── wal-000001.log
└── wal-000002.log
5.3 Chunk Write Strategies
// crate: rustfs-chunk-service
// file: crates/chunk-service/src/write.rs
impl ChunkServiceImpl {
/// Route write to the correct strategy based on storage policy.
pub async fn write_chunk(
&self,
req: WriteChunkReq,
policy: &StoragePolicy,
) -> Result<WriteChunkResp> {
match policy {
StoragePolicy::Replicated { .. } => {
self.write_replicated(req).await
}
StoragePolicy::ErasureCoded { data_shards, parity_shards, .. } => {
let block_size = EC_BLOCK_SIZE; // 1 MiB
let write_size = req.data.len() as u64;
if write_size >= block_size
&& req.offset % block_size == 0
{
// Full-block write: fast path, no RMW
self.write_ec_full_blocks(req, *data_shards, *parity_shards).await
} else {
// Partial write: read-modify-write
self.write_ec_partial(req, *data_shards, *parity_shards).await
}
}
_ => unreachable!(),
}
}
/// Replicated write: write to all replicas in parallel.
///
/// Latency = max(replica_write_latencies) — bounded by slowest disk.
/// For block volumes, this is the primary write path.
async fn write_replicated(&self, req: WriteChunkReq) -> Result<WriteChunkResp> {
let chunk_path = self.chunk_path(&req.chunk_id);
// Open chunk file (create if necessary)
let file = tokio::fs::OpenOptions::new()
.write(true)
.create(req.create)
.open(&chunk_path)
.await?;
// Write at offset
file.seek(SeekFrom::Start(req.offset)).await?;
file.write_all(&req.data).await?;
if req.sync {
file.sync_data().await?;
}
// Update bitrot checksum for the affected range
let checksum = self.update_chunk_checksum(
&req.chunk_id, req.offset, &req.data
).await?;
Ok(WriteChunkResp {
bytes_written: req.data.len() as u64,
checksum,
})
}
/// EC full-block write: encode data blocks → parity blocks, write all shards.
///
/// Reuses existing ecstore erasure_coding::encode path.
async fn write_ec_full_blocks(
&self,
req: WriteChunkReq,
data_shards: u32,
parity_shards: u32,
) -> Result<WriteChunkResp> {
// Split into EC block-aligned segments
// Encode each block → data + parity shards
// Write each shard to its designated disk
// (This is the existing ecstore SetDisks::put_object logic,
// refactored to operate on individual blocks rather than whole objects)
todo!("Reuse ecstore erasure_coding::encode")
}
}
6. Client Implementations
6.1 FUSE Client (rustfs-mount)
Binary: rustfs-mount
Usage: rustfs-mount --mds=mds1:9100,mds2:9100 --mountpoint=/mnt/rustfs [options]
Options:
--read-ahead=32m Max read-ahead window per file
--cache-size=4g Client-side data cache (page cache supplement)
--attr-timeout=5s Attribute cache TTL (overridden by leases)
--entry-timeout=5s Dentry cache TTL (overridden by leases)
--direct-io Default to O_DIRECT (bypass client cache)
--storage-policy=ec:4+2 Default storage policy for new files
--worker-threads=16 FUSE worker threads
--max-write=1m Max write size per FUSE request
6.1.1 FUSE Client Architecture
// crate: rustfs-fuse
// file: crates/fuse-client/src/client.rs
pub struct RustfsClient {
// Connections
mds: MdsClient,
dss_pool: DssConnectionPool,
// Caches
inode_cache: LruCache<InodeId, CachedInode>,
dentry_cache: LruCache<(InodeId, String), CachedDentry>,
chunk_map_cache: LruCache<InodeId, Arc<ChunkMap>>,
// State
open_files: DashMap<FileHandle, OpenFile>,
session: Session,
next_fh: AtomicU64,
// Engines
read_ahead: ReadAheadEngine,
lease_manager: ClientLeaseManager,
write_buffer: WriteBufferManager,
}
struct OpenFile {
ino: InodeId,
flags: u32,
chunk_map: Arc<ChunkMap>,
lease: Option<Lease>,
dirty: AtomicBool,
}
struct CachedInode {
attr: InodeAttr,
valid_until: Instant,
lease: Option<Lease>, // if lease held, cache is valid until revoked
}
6.1.2 FUSE Read Path (Hot Path for AI)
// crate: rustfs-fuse
// file: crates/fuse-client/src/read.rs
impl RustfsClient {
/// FUSE read handler. Called for every read() syscall.
///
/// Performance target: <5ms first-byte latency for cached chunk maps,
/// aggregate bandwidth scales linearly with DSS node count.
pub async fn fuse_read(
&self,
ino: InodeId,
fh: FileHandle,
offset: u64,
size: u32,
) -> Result<Vec<u8>> {
let file = self.open_files.get(&fh)
.ok_or(Error::BadFileHandle)?;
// Step 1: Get chunk map (cached if layout lease held)
let chunk_map = self.get_chunk_map_cached(ino).await?;
// Step 2: Resolve byte range → chunk reads
let chunk_reads = chunk_map.resolve_range(offset, size as u64);
// Step 3: Issue parallel reads to DSS nodes
//
// KEY OPTIMIZATION: each chunk maps to a different DSS node
// (due to striping), so all reads go in parallel to different
// servers. For a 64 MiB read with 4 MiB chunks = 16 parallel
// fetches to up to 16 different DSS nodes.
let mut fetches = Vec::with_capacity(chunk_reads.len());
for entry in &chunk_reads {
let loc = self.pick_best_location(&entry.locations);
let dss = self.dss_pool.get(loc.node_id);
fetches.push(dss.read_chunk(ReadChunkReq {
chunk_id: loc.chunk_id.clone(),
offset: entry.offset_in_chunk,
length: entry.length,
direct_io: file.flags & O_DIRECT != 0,
verify: true,
}));
}
let results = futures::future::try_join_all(fetches).await?;
// Step 4: Assemble results into contiguous buffer
let mut buf = Vec::with_capacity(size as usize);
for result in results {
buf.extend_from_slice(&result.data);
}
// Step 5: Trigger read-ahead
self.read_ahead.on_read(ino, offset, size as u64, &chunk_map);
Ok(buf)
}
/// Pick the best replica/shard location for a read.
///
/// Priority: local node > same rack > lowest latency > round-robin
fn pick_best_location(&self, locs: &[ChunkLocation]) -> &ChunkLocation {
// 1. Check if any location is on this client's local node
// (for DSS nodes that are also compute nodes)
if let Some(local) = locs.iter().find(|l| l.node_id == self.local_node_id) {
return local;
}
// 2. Pick by lowest recent latency
locs.iter()
.min_by_key(|l| self.dss_pool.latency_ns(l.node_id))
.unwrap()
}
}
6.1.3 Read-Ahead Engine
// crate: rustfs-fuse
// file: crates/fuse-client/src/readahead.rs
/// Adaptive read-ahead engine.
///
/// Detects sequential, strided, and random access patterns per-file.
/// For sequential access, exponentially grows prefetch window up to max.
/// For strided access (common in DataLoader: read shard, jump, read next),
/// learns the stride and prefetches accordingly.
///
/// Design: mirrors Linux kernel's readahead algorithm but runs in userspace
/// and operates at chunk granularity rather than page granularity.
pub struct ReadAheadEngine {
trackers: DashMap<InodeId, AccessTracker>,
prefetcher: mpsc::Sender<PrefetchCmd>,
config: ReadAheadConfig,
}
pub struct ReadAheadConfig {
pub max_window: u64, // bytes (default: 32 MiB)
pub initial_window: u64, // bytes (default: chunk_size)
pub growth_factor: f64, // default: 2.0
pub shrink_threshold: u32, // sequential break count to shrink (default: 2)
}
struct AccessTracker {
last_offset: u64,
last_size: u64,
window: u64,
sequential_n: u32,
stride: Option<u64>,
stride_hits: u32,
pattern: AccessPattern,
}
enum AccessPattern {
Unknown,
Sequential,
Strided { stride: u64 },
Random,
}
impl ReadAheadEngine {
/// Called after every successful read.
pub fn on_read(
&self,
ino: InodeId,
offset: u64,
size: u64,
chunk_map: &ChunkMap,
) {
let mut tracker = self.trackers.entry(ino).or_default();
// Classify access pattern
let expected_next = tracker.last_offset + tracker.last_size;
let is_sequential = offset == expected_next;
let is_strided = tracker.stride.map_or(false, |s| {
offset == tracker.last_offset + s
});
match (is_sequential, is_strided) {
(true, _) => {
tracker.pattern = AccessPattern::Sequential;
tracker.sequential_n += 1;
tracker.window = (tracker.window as f64 * self.config.growth_factor)
as u64;
tracker.window = tracker.window.min(self.config.max_window);
}
(false, true) => {
tracker.pattern = AccessPattern::Strided {
stride: tracker.stride.unwrap()
};
tracker.stride_hits += 1;
tracker.window = (tracker.window as f64 * self.config.growth_factor)
as u64;
tracker.window = tracker.window.min(self.config.max_window);
}
(false, false) => {
// Try to learn a new stride
let gap = offset.wrapping_sub(tracker.last_offset);
if gap > 0 && gap < chunk_map.chunk_size * 1024 {
tracker.stride = Some(gap);
tracker.stride_hits = 0;
}
tracker.sequential_n = 0;
tracker.window = self.config.initial_window;
}
}
// Issue prefetch
if tracker.window > 0 {
let prefetch_start = match tracker.pattern {
AccessPattern::Sequential => offset + size,
AccessPattern::Strided { stride } => offset + stride,
_ => return,
};
let chunks = chunk_map.resolve_range(prefetch_start, tracker.window);
let _ = self.prefetcher.try_send(PrefetchCmd {
ino,
chunks,
priority: tracker.sequential_n,
});
}
tracker.last_offset = offset;
tracker.last_size = size;
}
}
6.2 Block Device Exports
6.2.1 iSCSI Target
// crate: rustfs-iscsi
// file: crates/iscsi-target/src/target.rs
/// iSCSI target that exports RustFS block volumes as SCSI LUNs.
///
/// Each volume maps to one LUN on the target.
/// Supports multi-path: volume can be exported from multiple targets.
///
/// Target naming: iqn.2026-06.com.rustfs:{volume_name}
pub struct RustfsIscsiTarget {
mds: MdsClient,
dss_pool: DssConnectionPool,
volumes: DashMap<Lun, AttachedVolume>,
wal: WriteAheadLog,
}
struct AttachedVolume {
volume: Volume,
chunk_map: ChunkMap, // cached from MDS
dirty_map: RwLock<BTreeSet<u64>>, // dirty chunk indices
}
impl ScsiTarget for RustfsIscsiTarget {
async fn execute_command(&self, cmd: ScsiCommand) -> ScsiResult {
match cmd.opcode {
SCSI_READ_16 => {
let lba = cmd.lba();
let blocks = cmd.transfer_length();
let vol = self.volumes.get(&cmd.lun)?;
let byte_offset = lba * vol.volume.block_size as u64;
let byte_length = blocks as u64 * vol.volume.block_size as u64;
// Resolve chunks and parallel read
let data = self.read_volume_range(
&vol, byte_offset, byte_length
).await?;
ScsiResult::data(data)
}
SCSI_WRITE_16 => {
let lba = cmd.lba();
let vol = self.volumes.get(&cmd.lun)?;
let byte_offset = lba * vol.volume.block_size as u64;
// WAL entry for write ordering
self.wal.append(WalEntry::Write {
volume_id: vol.volume.id.clone(),
offset: byte_offset,
data: cmd.data().to_vec(),
}).await?;
// Write to chunk storage
self.write_volume_range(
&vol, byte_offset, cmd.data()
).await?;
ScsiResult::ok()
}
SCSI_SYNCHRONIZE_CACHE => {
let vol = self.volumes.get(&cmd.lun)?;
self.flush_volume(&vol).await?;
ScsiResult::ok()
}
SCSI_INQUIRY => {
ScsiResult::inquiry(InquiryData {
vendor: b"RustFS ",
product: b"Block Volume ",
revision: b"0100",
})
}
SCSI_READ_CAPACITY_16 => {
let vol = self.volumes.get(&cmd.lun)?;
let blocks = vol.volume.capacity / vol.volume.block_size as u64;
ScsiResult::capacity(blocks - 1, vol.volume.block_size)
}
_ => ScsiResult::check_condition(ILLEGAL_REQUEST),
}
}
}
6.2.2 NVMe-oF Target (High-Performance Path)
Binary: rustfs-nvmeof
Usage: rustfs-nvmeof --mds=mds1:9100 --transport=tcp --port=4420
Exposes volumes as NVMe namespaces.
For RDMA transport, use --transport=rdma (requires RDMA NIC).
Implementation: wraps SPDK NVMe-oF target library via FFI,
or uses kernel nvmet for TCP transport.
Maps:
NVMe Read → ReadChunk (parallel per chunk)
NVMe Write → WriteChunk (replicated)
NVMe Flush → FlushChunk
NVMe TRIM → TruncateChunk (deallocate)
7. Performance Design
7.1 AI Read Bandwidth Target
Goal: saturate GPU cluster's storage bandwidth
Cluster: 128 GPU nodes, each with 100 Gbps NIC
Aggregate demand: 128 × 12.5 GB/s = 1.6 TB/s
DSS cluster: 32 nodes, each with 4 × NVMe (3 GB/s each) + 100 Gbps NIC
Aggregate supply: 32 × 12 GB/s = 384 GB/s (disk-limited)
32 × 12.5 GB/s = 400 GB/s (network-limited)
Effective: ~380 GB/s
Gap: 1.6 TB/s demand vs 380 GB/s supply → 4.2x shortfall
Solution: client-side caching + read-ahead
Training epochs re-read the same data. After epoch 1:
- 90%+ reads served from client page cache (ReadData lease held)
- Only epoch 1 needs DSS bandwidth
- Client needs enough RAM to cache its shard of the dataset
With 128 GPU nodes, each caching its 1/128 fraction:
Dataset 10 TB → each node caches ~80 GB → fits in 128-256 GB RAM
Epoch 1: 10 TB / 380 GB/s = ~26 seconds
Epoch 2+: served from client cache → limited by memory bandwidth (~100+ GB/s per node)
7.2 Latency Targets
Operation Target Notes
───────────────────────────── ────────── ──────────────────────────
POSIX stat (cached) < 0.01 ms Served from client inode cache
POSIX stat (uncached) < 1 ms 1 MDS gRPC round trip
POSIX open < 2 ms MDS lookup + chunk map fetch
POSIX read (cached chunk) < 0.1 ms Client page cache hit
POSIX read (1 chunk, uncached)< 5 ms 1 DSS gRPC round trip
POSIX readdir (100 entries) < 2 ms 1 MDS gRPC round trip
POSIX create < 3 ms MDS create + lease grant
Block read 4K (cached) < 0.1 ms iSCSI initiator cache
Block read 4K (uncached) < 2 ms 1 DSS round trip
Block write 4K < 1 ms Write to nearest replica + WAL
Block flush < 5 ms Sync all dirty chunks
S3 GET (small object) < 10 ms MDS lookup + 1 DSS read
S3 PUT (small object) < 15 ms MDS create + 1 DSS write + commit
S3 LIST (1000 objects) < 5 ms MDS readdir (O(1) not O(N)!)
7.3 Client-Side Caching Hierarchy
┌─────────────────────────────────────────────────┐
│ FUSE CLIENT CACHES │
│ │
│ ┌──────────────────────────────────────────┐ │
│ │ Level 1: Linux VFS Page Cache │ │
│ │ (kernel manages, FUSE honors via TTL) │ │
│ │ Size: controlled by OS (typically 50% │ │
│ │ of free RAM) │ │
│ │ Eviction: LRU │ │
│ │ Coherency: invalidated on lease revoke │ │
│ └──────────────────────────────────────────┘ │
│ │
│ ┌──────────────────────────────────────────┐ │
│ │ Level 2: Client Chunk Cache │ │
│ │ (userspace, managed by rustfs-mount) │ │
│ │ Size: --cache-size (default 4 GB) │ │
│ │ Eviction: LRU with frequency boost │ │
│ │ Coherency: tied to ReadData lease │ │
│ └──────────────────────────────────────────┘ │
│ │
│ ┌──────────────────────────────────────────┐ │
│ │ Level 3: Inode + Dentry + ChunkMap Cache│ │
│ │ (metadata caches in rustfs-mount) │ │
│ │ Size: 100K entries (inode), 500K (dentry│) │
│ │ Coherency: tied to ReadMeta/Layout lease│ │
│ └──────────────────────────────────────────┘ │
│ │
└─────────────────────────────────────────────────┘
8. Crate Map and Build Targets
8.1 New Crates
crates/
├── inode/ # 3K lines — core types (InodeAttr, ChunkMap, Volume, Lease)
│ ├── src/
│ │ ├── types.rs # InodeId, InodeAttr, InodeKind, StoragePolicy
│ │ ├── dentry.rs # DirEntry, DirListing
│ │ ├── chunkmap.rs # ChunkMap, ChunkLocation, resolve_range()
│ │ ├── volume.rs # Volume, VolumeSnapshot, ExtentMap
│ │ ├── lease.rs # Lease, LeaseKind, ByteRangeLock
│ │ └── lib.rs
│ └── Cargo.toml # deps: serde, rmp-serde
│
├── mds/ # 18K lines — metadata server
│ ├── src/
│ │ ├── service.rs # gRPC service implementation (NamespaceService, etc.)
│ │ ├── store.rs # MetadataStore trait
│ │ ├── rocksdb.rs # RocksDB backend
│ │ ├── raft.rs # Raft consensus wrapper (openraft)
│ │ ├── lease_mgr.rs # Server-side lease manager + revocation
│ │ ├── lock_mgr.rs # Byte-range lock manager
│ │ ├── alloc.rs # Chunk allocation policy (balancing across DSS nodes)
│ │ ├── s3bridge.rs # S3 path ↔ inode translation
│ │ ├── gc.rs # Garbage collection for orphaned chunks
│ │ └── main.rs # MDS binary entry point
│ └── Cargo.toml # deps: tonic, rocksdb, openraft, rustfs-inode
│
├── chunk-service/ # 6K lines — DSS chunk I/O service
│ ├── src/
│ │ ├── service.rs # gRPC ChunkService implementation
│ │ ├── write.rs # Write strategies (replicated, EC full, EC partial)
│ │ ├── read.rs # Read with bitrot verify
│ │ ├── layout.rs # On-disk chunk file management
│ │ └── lib.rs
│ └── Cargo.toml # deps: tonic, rustfs-io-core, rustfs-ecstore-core
│
├── fuse-client/ # 12K lines — FUSE client
│ ├── src/
│ │ ├── client.rs # RustfsClient struct
│ │ ├── fuse_ops.rs # fuser::Filesystem implementation
│ │ ├── read.rs # Read path with parallel chunk fetch
│ │ ├── write.rs # Write path with buffering + commit
│ │ ├── readahead.rs # Adaptive read-ahead engine
│ │ ├── cache.rs # Client-side chunk and metadata caches
│ │ ├── lease.rs # Client-side lease manager + revocation handler
│ │ └── main.rs # rustfs-mount binary entry point
│ └── Cargo.toml # deps: fuser, tonic, rustfs-inode
│
├── iscsi-target/ # 8K lines — iSCSI target
│ ├── src/
│ │ ├── target.rs # SCSI command handler
│ │ ├── session.rs # iSCSI session/connection management
│ │ ├── wal.rs # Write-ahead log for block write ordering
│ │ └── main.rs # rustfs-iscsi binary entry point
│ └── Cargo.toml
│
└── nvmeof-target/ # 5K lines — NVMe-oF target (phase 2)
├── src/
│ ├── target.rs # NVMe command handler
│ └── main.rs
└── Cargo.toml
8.2 Modified Existing Crates
crates/ecstore/ # DECOMPOSE into:
→ ecstore-core/ # EC encode/decode, quorum, shard math
→ ecstore-disk/ # LocalDisk, disk health
→ ecstore-pool/ # EndpointServerPools, set management
→ ecstore-s3/ # S3 object operations (existing SetDisks logic)
crates/lock/ # EXTEND:
+ posix_lock.rs # ByteRangeLock support
+ lease.rs # Server-side lease tracking (used by MDS)
crates/filemeta/ # EXTEND:
+ posix_inode_id field in ObjectInfo (bridge to unified namespace)
crates/protos/ # ADD:
+ mds.proto # MDS service definitions
+ chunk.proto # DSS chunk service definitions
rustfs/src/storage/ # ADD:
+ unified.rs # UnifiedObjectLayer (S3 API → MDS + DSS)
8.3 Build Targets
Binary Crate Entry Description
─────────────────── ──────────────────── ─────────────────────────────
rustfs rustfs/src/main.rs S3 API + admin console (existing)
+ unified namespace mode (new)
rustfs-mds crates/mds/main.rs Metadata server
rustfs-mount crates/fuse-client/ FUSE client (POSIX mount)
main.rs
rustfs-iscsi crates/iscsi-target/ iSCSI target for block volumes
main.rs
rustfs-nvmeof crates/nvmeof- NVMe-oF target (future)
target/main.rs
9. Configuration
9.1 Cluster Configuration
# /etc/rustfs/cluster.toml
[cluster]
id = "prod-ai-cluster-01"
mode = "unified" # "object-only" | "unified"
[mds]
endpoints = ["mds1:9100", "mds2:9100", "mds3:9100"]
replication = "raft" # "none" | "raft"
[dss]
# DSS nodes auto-register with MDS at startup.
# This section defines defaults.
chunk_size = "4MiB"
default_storage_policy = "ec:4+2"
[namespace]
# S3 buckets appear under this path in the POSIX namespace
bucket_root = "/buckets"
# Block volumes appear under this path
volume_root = "/volumes"
# Default path for POSIX files not associated with an S3 bucket
default_root = "/"
[leases]
heartbeat_interval = "30s"
lease_timeout = "90s" # revoke if no heartbeat for this long
read_lease_ttl = "300s" # max lease duration before forced renewal
[readahead]
max_window = "32MiB"
initial_window = "4MiB"
[block]
default_replica_factor = 3
wal_sync_mode = "fdatasync" # "fdatasync" | "fsync" | "none"
thin_provision = true
9.2 Storage Policy Profiles
# /etc/rustfs/policies.toml
[profiles.ai-training]
# Optimized for write-once, read-many from hundreds of GPU nodes.
storage_policy = "ec:8+4" # high durability, max throughput
chunk_size = "4MiB" # matches typical DataLoader read size
default_lease = "ReadData" # auto-grant read leases
readahead_max = "64MiB" # aggressive prefetch
[profiles.ai-checkpoints]
# Optimized for periodic large sequential writes + occasional reads.
storage_policy = "ec:4+2"
chunk_size = "4MiB"
default_lease = "WriteData"
[profiles.vm-block]
# Optimized for random 4K I/O.
storage_policy = "replicated:3"
chunk_size = "4MiB"
default_lease = "none" # block volumes manage coherency via iSCSI
wal_enabled = true
[profiles.scratch]
# General purpose POSIX workspace.
storage_policy = "replicated:2" # lower durability, faster writes
chunk_size = "1MiB"
10. Implementation Phases
Phase 1 — Foundation (Months 1-3)
Deliverables:
✓ crates/inode/ — all types compiled and tested
✓ crates/mds/ — single-node RocksDB MDS with namespace + chunk map APIs
✓ crates/chunk-service/ — replicated chunk read/write on DSS nodes
✓ crates/fuse-client/ — basic FUSE mount with open/read/write/stat/readdir/mkdir
Milestone: mount /mnt/rustfs, create files, read them back, run ls -la
Acceptance test:
$ rustfs-mds --standalone &
$ rustfs-mount --mds=localhost:9100 /mnt/rustfs &
$ dd if=/dev/urandom of=/mnt/rustfs/test.bin bs=4M count=256 # 1 GB write
$ md5sum /mnt/rustfs/test.bin # read back
$ ls -la /mnt/rustfs/ # readdir
Phase 2 — AI-Ready (Months 3-5)
Deliverables:
✓ Read-ahead engine with sequential + stride detection
✓ ReadData lease for multi-reader caching
✓ Layout lease for chunk map caching
✓ EC-striped chunks (reuse ecstore encode/decode)
✓ Parallel multi-DSS chunk fetch in FUSE read path
✓ S3 API routed through MDS (UnifiedObjectLayer)
Milestone: train a real model using POSIX mount
Acceptance test:
$ torchrun --nproc_per_node=8 train.py \
--data_path=/mnt/rustfs/buckets/training-data/imagenet/
# 8 GPU processes reading concurrently via FUSE
# Verify: epoch 2 is faster than epoch 1 (cache warm)
# Verify: S3 PUT to "training-data" bucket visible in /mnt/rustfs/
Phase 3 — Block Storage (Months 5-7)
Deliverables:
✓ VolumeService in MDS
✓ Volume extent map with thin provisioning
✓ COW snapshots
✓ WAL for write ordering on DSS
✓ crates/iscsi-target/ — iSCSI target daemon
Milestone: boot a VM from RustFS block volume
Acceptance test:
$ rustfs-admin volume create --name=vm-boot --size=50G --thin
$ rustfs-iscsi --mds=mds1:9100 --volume=vm-boot --lun=0 &
$ qemu-system-x86_64 \
-drive file=iscsi://localhost/iqn.2026-06.com.rustfs:vm-boot/0 \
-m 4G
# VM boots from RustFS block volume
# Verify: snapshot, clone, resize operations
Phase 4 — Production (Months 7-10)
Deliverables:
✓ HA MDS with Raft (3-node or 5-node cluster)
✓ MDS sharding (subtree partitioning)
✓ ecstore decomposition (monolith → sub-crates)
✓ Kubernetes CSI driver (RWX for POSIX, RWO for block)
✓ NVMe-oF target for high-performance block
Milestone: pass production readiness review
Phase 5 — Performance (Months 10-14)
Deliverables:
✓ RDMA transport for chunk reads (DSS ↔ client)
✓ GPUDirect Storage integration (bypass CPU on reads)
✓ io_uring FUSE passthrough (bypass FUSE kernel overhead)
✓ Client write-back caching with lease coherency
✓ Automated tiering (hot replicated → cold EC)
Milestone: match Lustre/BeeGFS on IOR and mdtest benchmarks
11. Compatibility Matrix
Feature S3 API POSIX Block
─────────────────────── ────────────── ───────────── ─────────────
Create file/object PUT Object creat()/open() create-vol
Read GET Object read() SCSI READ
Write PUT Object write() SCSI WRITE
Delete DELETE Object unlink() delete-vol
List ListObjects readdir() list-vols
Rename COPY+DELETE rename() N/A
Permissions Bucket Policy chmod/chown LUN masking
+ IAM + POSIX ACL
Versioning S3 Versioning (via snapshots) COW snapshots
Locking S3 Object Lock flock/fcntl iSCSI reservations
Notifications S3 Events inotify (future) N/A
Encryption SSE-S3/SSE-KMS (inherited) (inherited)
Compression (inherited) (inherited) (disabled)
Quotas Bucket quota Dir quota Vol capacity
Replication Bucket repl. (inherited) Vol repl factor
12. Open Design Questions
ID Question Decision needed by
──── ──────────────────────────────────────────── ─────────────────
DQ1 Should MDS store chunk maps inline with Phase 1 start
inodes (simpler) or in a separate CF
(better for large files with millions of
chunks)?
DQ2 For the S3 ↔ POSIX bridge, should we Phase 2 start
support hard links between S3 "objects"?
(S3 has no link concept)
DQ3 For block volumes, should we use a separate Phase 3 start
WAL per volume or a shared WAL per DSS node?
DQ4 RDMA transport: use existing tonic/gRPC Phase 5 start
with RDMA transport, or implement custom
zero-copy RDMA verbs?
DQ5 Should the FUSE client support direct Phase 2 start
kernel bypass (io_uring passthrough or
CUSE) from day 1, or add it later?
DQ6 MDS sharding strategy: static subtree Phase 4 start
assignment or dynamic migration (like
CephFS)?
Appendix A: Reference Architectures
System MDS Data Path Block
─────────── ─────────────────── ────────────────── ────────────
Lustre Separate MDS (MDD OST servers, file N/A
on ldiskfs/ZFS) striping, LDLM
BeeGFS Separate MDS OST servers, N/A (BeeGFS
(BuddyMirror HA) file striping Mon only)
CephFS MDS cluster RADOS objects Ceph RBD
(Raft-like) (EC or replicated) (same RADOS)
JuiceFS External DB Any S3-compatible N/A
(Redis/TiKV/MySQL) object store
RustFS Embedded in Erasure-coded N/A
(current) ecstore (no shards on local
separate MDS) disks
RustFS Separate MDS Chunk service on Volume manager
(proposed) (RocksDB + Raft) existing DSS nodes + iSCSI/NVMe-oF
(EC or replicated) on same DSS
End of specification. This document should be kept in sync with implementation as the codebase evolves. Update the version number and revision date on every significant change.