RustFS 文档 文档

RustFS Technical Review: Feasibility for POSIX Parallel FS & Block Storage

RustFS Technical Review: Feasibility for POSIX Parallel FS & Block Storage

Executive Summary

RustFS is an S3-compatible object storage system — not a POSIX filesystem and not a block device. It cannot currently serve as a BeeGFS/Lustre replacement for AI training workloads that use fseek()/fread(), nor as a VM/bare-metal block disk. Achieving those goals requires building several entirely new subsystems on top of or alongside the existing codebase.

This document reviews the architecture, identifies every gap, and proposes a concrete roadmap.


1. What RustFS Is Today

AspectCurrent State
Storage modelS3-compatible object storage (PUT/GET/DELETE whole objects)
Data pathHTTP → erasure coding → local disk I/O via ecstore (87K lines)
Protocol supportS3 API, OpenStack Swift, FTP/FTPS, WebDAV
Data durabilityErasure coding with bitrot protection and healing
ClusteringDistributed mode via gRPC inter-node RPC (still under testing)
Language98% Rust, Apache 2.0 license
MaturityBeta (1.0.0-beta.7), GA planned July 2026
POSIX supportNone
Block device supportNone
RDMA/DPUAnnounced on roadmap, separate gpu-cache repo exists (early stage)

2. What BeeGFS / Lustre Provide (The Target)

2.1 POSIX Parallel Filesystem Capabilities

AI training workloads (PyTorch DataLoader, TensorFlow tf.data, HuggingFace datasets) rely on:

fd = open("/mnt/cluster/dataset/shard-00042.bin", O_RDONLY);
fseek(fd, offset, SEEK_SET);     // random seek into middle of file
fread(buf, chunk_size, 1, fd);   // read a specific region
// ... repeat from many processes simultaneously

What BeeGFS/Lustre deliver that RustFS does not:

  • POSIX VFS integration: mount as a kernel filesystem or FUSE, exposing open/read/write/seek/close/stat/readdir/mmap
  • Byte-range I/O: read/write arbitrary byte ranges without downloading whole objects
  • File striping: a single file is split across N storage targets; parallel I/O from all targets simultaneously
  • Client-side caching / read-ahead: predictive prefetch for sequential and strided access patterns
  • Distributed metadata: separate MDS (Metadata Server) from OSS (Object Storage Server)
  • File locking: flock() / fcntl() advisory and mandatory locks
  • Concurrent readers: hundreds of GPU nodes reading different offsets of the same file simultaneously with near-linear aggregate bandwidth scaling

2.2 Block Device Capabilities (VM / Bare-Metal Disks)

For VMs (KVM/QEMU, VMware, Hyper-V) and bare-metal iSCSI/NVMe-oF boot:

  • Block-level I/O: fixed-size block reads/writes (typically 4K-64K), not object-level
  • Thin provisioning: allocate-on-write with overcommit
  • Snapshots & clones: instant copy-on-write snapshots
  • Consistency: write-ordering guarantees, flush/barrier support
  • Export protocols: iSCSI, NVMe-oF (NVMe over Fabrics), virtio-blk, NBD
  • Examples: Ceph RBD, Longhorn, OpenEBS, LVM on shared storage

3. Gap Analysis: RustFS vs. Requirements

3.1 POSIX Filesystem Layer — MISSING ENTIRELY

RequirementRustFS StatusGap Severity
Kernel VFS / FUSE mount pointNot implementedCRITICAL
open/read/write/seek/close syscallsObjects are HTTP GET/PUT onlyCRITICAL
Byte-range read (fseek + fread)S3 Range GET exists but no POSIX mappingHIGH
Byte-range write (random write into file)Not supported (objects are immutable-on-PUT)CRITICAL
File striping across nodesErasure coding != striping for parallel read throughputHIGH
Distributed metadata serverNo separate MDS; metadata is in ecstore per-bucketHIGH
readdir / stat / chmod / POSIX permissionsNot implementedCRITICAL
Client-side read-ahead / cachingNo client component existsHIGH
mmap() supportNot possible without VFS layerCRITICAL
File locking (flock / fcntl)Not implementedMEDIUM

3.2 Block Device Layer — MISSING ENTIRELY

RequirementRustFS StatusGap Severity
Fixed-block I/O (4K aligned R/W)Object-level onlyCRITICAL
iSCSI / NVMe-oF targetNot implementedCRITICAL
Thin provisioningNot applicable to object modelCRITICAL
Snapshots & clones (COW)S3 versioning exists, but no block-level COWHIGH
Write barriers / flush semanticsHTTP has no equivalentCRITICAL
NBD (Network Block Device) serverNot implementedCRITICAL

3.3 What RustFS Has That Can Be Reused

ComponentReusable For
ecstore erasure coding engineData durability for both POSIX and block layers
io-core zero-copy I/O + buffer poolFoundational I/O primitives
rio reader pipeline (encrypt → compress → hash)Data pipeline for all storage paths
gRPC inter-node RPC (protos/)Cluster communication backbone
IAM / policy engineMulti-tenant access control
lock distributed lock managerFoundation for file locking
Observability stack (Prometheus, OTel)Monitoring for new subsystems
heal + scannerBackground integrity checking

4. Proposed Architecture to Close the Gaps

4.1 Overall System Design

┌──────────────────────────────────────────────────────────┐
│                    CLIENT NODES                           │
│  ┌──────────────┐  ┌──────────────┐  ┌──────────────┐   │
│  │ FUSE Client  │  │ Kernel Client│  │ NBD / iSCSI  │   │
│  │ (rustfs-fuse)│  │ (rustfs.ko)  │  │   Client     │   │
│  └──────┬───────┘  └──────┬───────┘  └──────┬───────┘   │
│         │ POSIX syscalls   │ VFS ops         │ Block I/O │
└─────────┼──────────────────┼─────────────────┼───────────┘
          │                  │                 │
          ▼                  ▼                 ▼
┌──────────────────────────────────────────────────────────┐
│                   NETWORK LAYER                           │
│         gRPC / RDMA (future) / TCP                       │
└────────────┬───────────────────────┬─────────────────────┘
             │                       │
             ▼                       ▼
┌────────────────────┐    ┌─────────────────────┐
│   METADATA SERVER  │    │  STORAGE SERVERS     │
│   (rustfs-mds)     │    │  (rustfs-oss)        │
│                    │    │                      │
│ • Inode table      │    │ • Chunk storage      │
│ • Directory tree   │    │ • Erasure coding     │
│ • File→chunk map   │    │ • Block volumes      │
│ • POSIX attrs      │    │ • Replication         │
│ • Lock manager     │    │                      │
│ • Lease manager    │    │ (reuses ecstore,     │
│                    │    │  io-core, rio)       │
└────────────────────┘    └─────────────────────┘

4.2 POSIX Layer: What to Build

A. FUSE Client (rustfs-fuse) — Start Here

This is the fastest path to a working POSIX mount:

New crate: crates/fuse-client/
Dependencies: fuser (Rust FUSE library), gRPC client

Must implement:
├── init()          → connect to MDS + discover OSS nodes
├── lookup()        → MDS inode lookup
├── getattr()       → MDS stat
├── readdir()       → MDS directory listing
├── open()          → MDS: allocate file handle, get chunk map
├── read()          → OSS: parallel byte-range reads from striped chunks
├── write()         → OSS: write to chunk targets, update MDS
├── seek()          → client-local offset tracking
├── flush()/fsync() → OSS: ensure durability
├── release()       → MDS: release file handle + leases
├── mkdir/rmdir/unlink/rename → MDS operations
└── setattr/chmod/chown → MDS POSIX attribute updates

Key design decisions for AI workloads:

  1. Stripe size: 1MB–4MB chunks (matches typical AI data loader read sizes)
  2. Read-ahead: adaptive prefetch based on sequential detection (like Lustre's llite read-ahead)
  3. Client-side page cache: honor Linux VFS page cache for repeated reads
  4. Parallel chunk fetch: for a single read(fd, buf, 64MB), fetch all stripe chunks in parallel from different OSS nodes

B. Metadata Server (rustfs-mds)

New crate: crates/mds/

Core data structures:
├── InodeTable        → inode_id → {type, size, uid, gid, mode, timestamps, xattrs}
├── DirectoryTree     → parent_inode + name → child_inode
├── ChunkMap          → inode_id + offset → [(oss_node, chunk_id, stripe_idx)]
├── LockTable         → inode_id → [lock_range, lock_type, owner]
├── LeaseManager      → inode_id → [client_id, lease_type, expiry]
└── JournalLog        → WAL for crash recovery

Storage backend options:
  • RocksDB/sled for embedded single-MDS
  • Raft consensus (via openraft) for HA MDS cluster

Metadata scalability (following Lustre's DNE model): partition the namespace into subtrees, each managed by a different MDS instance. This allows metadata operations to scale horizontally.

C. Object Storage Server Adaptation (rustfs-oss)

The existing ecstore needs a new interface alongside S3:

New trait: ChunkStorageService (gRPC)

rpc WriteChunk(chunk_id, offset, data) → Result
rpc ReadChunk(chunk_id, offset, length) → Stream<bytes>
rpc DeleteChunks(chunk_ids) → Result
rpc ChunkStat(chunk_id) → {size, checksum, location}

This reuses the existing ecstore erasure engine, rio pipeline, and io-core zero-copy I/O — but exposes chunk-level (not object-level) operations.

4.3 Block Device Layer: What to Build

A. Volume Manager (rustfs-volume)

New crate: crates/volume/

Manages block volumes as sequences of fixed-size extents:

Volume:
  ├── volume_id: UUID
  ├── size: u64 (logical size)
  ├── block_size: 4096
  ├── extent_size: 4MB
  ├── extent_map: BTreeMap<extent_offset, (oss_node, chunk_id)>
  ├── snapshot_tree: COW B-tree
  └── provisioning: Thin | Thick

Operations:
  ├── create_volume(size, thin?) → volume_id
  ├── read_block(volume_id, lba, count) → data
  ├── write_block(volume_id, lba, data) → Result
  ├── flush(volume_id) → Result  (write barriers)
  ├── snapshot(volume_id) → snapshot_id
  ├── clone(snapshot_id) → new_volume_id
  └── delete_volume(volume_id) → Result

Thin provisioning: allocate extents on first write. The extent map starts empty; writes allocate new chunks on OSS nodes.

Copy-on-write snapshots: when creating a snapshot, freeze the extent map. Subsequent writes to the parent COW the modified extents to new locations.

B. Block Export Targets

iSCSI Target (widest compatibility — works with VMware, KVM, Windows, bare-metal):

New crate: crates/iscsi-target/
Dependency: existing Rust iSCSI target libraries or custom implementation

Maps:
  SCSI READ(10/16)  → volume.read_block()
  SCSI WRITE(10/16) → volume.write_block()
  SYNCHRONIZE_CACHE → volume.flush()
  INQUIRY/READ_CAPACITY → volume metadata

NVMe-oF Target (highest performance — for RDMA-capable environments):

New crate: crates/nvmeof-target/
Requires: SPDK bindings or kernel NVMe target integration

Maps NVMe commands to volume operations.
Best combined with RDMA for lowest latency.

NBD Server (simplest — useful for Linux VMs and testing):

New crate: crates/nbd-server/
Simplest protocol: NBD_CMD_READ/WRITE/FLUSH/TRIM → volume ops

4.4 RDMA Integration (Future, Aligns with RustFS Roadmap)

RustFS already plans RDMA/DPU support. For the POSIX and block layers:

Replace gRPC chunk transfer with:
├── RDMA READ  → zero-copy chunk fetch from OSS memory to client memory
├── RDMA WRITE → zero-copy chunk write from client to OSS
└── Benefits: bypass kernel network stack, ~1µs latency, 200+ Gbps

For the block layer specifically, RDMA + NVMe-oF gives near-local-disk latency, which is essential for VM boot disks.


5. AI Workload Compatibility: The "One Write, Multi fseek/fread" Pattern

This is the dominant pattern in AI training:

# Training data preparation (WRITE ONCE)
with open("/mnt/rustfs/dataset.bin", "wb") as f:
    for shard in shards:
        f.write(shard)  # sequential write, striped across OSS nodes

# Training (MULTI-PROCESS RANDOM READ)
# 8 GPUs × 4 DataLoader workers = 32 concurrent readers
def worker(rank):
    fd = open("/mnt/rustfs/dataset.bin", "rb")
    for batch_idx in assigned_batches(rank):
        fd.seek(batch_idx * batch_bytes)   # random seek
        data = fd.read(batch_bytes)        # read batch
        yield process(data)

Required optimizations for this pattern:

  1. Large stripe size (4MB): matches typical batch read sizes, avoids excessive metadata lookups
  2. Parallel stripe fetch: a single read() spanning multiple stripes fetches from all OSS nodes simultaneously
  3. Client read-ahead: detect sequential-within-stride patterns and prefetch next chunks
  4. Shared read leases: MDS grants shared-read leases to all clients on read-only files (no lock contention)
  5. Page cache cooperation: let Linux VFS cache recently-read chunks; critical for epoch-over-epoch data reuse
  6. Direct I/O option: for datasets that exceed RAM, bypass page cache and read straight from OSS via O_DIRECT

6. Implementation Roadmap

Phase 1: POSIX Foundation (3–4 months)

  • Design and implement MDS with inode table, directory tree, chunk map
  • Build rustfs-fuse client with basic open/read/write/seek/close/readdir/stat
  • Add chunk-level gRPC service to existing storage servers
  • File striping: write stripes across N storage targets
  • Basic read: parallel chunk fetch for single read() call
  • Milestone: mount filesystem, cp a file, cat it back

Phase 2: AI-Ready POSIX (2–3 months)

  • Client-side read-ahead and sequential detection
  • Shared read leases for concurrent readers
  • Large-file optimizations (adaptive stripe size, prefetch tuning)
  • O_DIRECT support for GPU-direct data loading
  • PyTorch DataLoader integration testing
  • Milestone: train a model on 100+ GPU nodes reading from RustFS

Phase 3: Block Storage (3–4 months)

  • Volume manager with extent mapping and thin provisioning
  • COW snapshot engine
  • NBD server (simplest export, for initial testing)
  • iSCSI target (production VMs)
  • Milestone: boot a VM from a RustFS block volume

Phase 4: Performance & Production (2–3 months)

  • RDMA data path for chunk transfer (client ↔ OSS)
  • NVMe-oF target for high-performance block export
  • Kernel FUSE bypass (io_uring FUSE or native kernel module) for lower latency
  • HA MDS with Raft consensus
  • Distributed namespace (DNE-style metadata partitioning)
  • Milestone: match BeeGFS throughput on standard benchmarks (IOR, mdtest)

Phase 5: Enterprise Features (ongoing)

  • Unified namespace (S3 + POSIX + Block see same data where applicable)
  • Quota management per-user / per-directory
  • HSM (Hierarchical Storage Management): auto-tier cold data
  • GPUDirect Storage integration (bypass CPU entirely)
  • Kubernetes CSI driver (ReadWriteMany for POSIX, ReadWriteOnce for block)

7. Comparable Projects for Reference

ProjectApproachLesson for RustFS
LustreKernel client + MDS + OSS, file stripingGold standard for HPC POSIX; study client read-ahead design
BeeGFSUserspace client (FUSE optional), MDS + OSSSimpler than Lustre; good reference for metadata design
CephFSFUSE + kernel client on top of RADOS object storeProves object store → POSIX is viable; study MDS design
Ceph RBDBlock device on top of RADOSProves object store → block is viable; study COW snapshots
JuiceFSFUSE client + any object store backend + Redis/TiKV MDSClosest model — uses S3 as backend, adds POSIX via FUSE
SeaweedFSObject store + FUSE mount + volume serverShows the object-to-POSIX bridge pattern

JuiceFS is the most directly relevant reference — it adds a POSIX layer on top of S3-compatible storage with a separate metadata engine. RustFS could follow a similar architecture but with tighter integration since it owns the storage engine.


8. Verdict

RustFS has strong foundational components (erasure coding, zero-copy I/O, Rust safety, gRPC cluster communication) that can be reused. But it is fundamentally an object store, and the two target use cases — POSIX parallel filesystem and block storage — require building multiple new subsystems that don't exist today:

What's NeededEstimated New CodeDifficulty
Metadata server (MDS)~15-25K linesHard
FUSE client~10-15K linesMedium-Hard
Chunk-level gRPC service~3-5K linesMedium
Volume manager + COW~10-15K linesHard
iSCSI/NVMe-oF/NBD targets~8-12K lines eachMedium-Hard
RDMA transport~5-10K linesHard

Total new code: roughly 50-80K lines of Rust, comparable in size to the existing ecstore crate alone. This is a 12-18 month engineering effort for a team of 4-6 experienced systems programmers.

The alternative — and possibly faster path — is to use JuiceFS or a similar POSIX gateway in front of RustFS's S3 API for the POSIX use case, and integrate an existing block-storage layer (like Longhorn or a custom thin-provisioning layer) for block devices. This sacrifices tight integration but gets to production much faster.