RustFS Technical Review: Feasibility for POSIX Parallel FS & Block Storage
RustFS Technical Review: Feasibility for POSIX Parallel FS & Block Storage
Executive Summary
RustFS is an S3-compatible object storage system — not a POSIX filesystem and not a block device. It cannot currently serve as a BeeGFS/Lustre replacement for AI training workloads that use fseek()/fread(), nor as a VM/bare-metal block disk. Achieving those goals requires building several entirely new subsystems on top of or alongside the existing codebase.
This document reviews the architecture, identifies every gap, and proposes a concrete roadmap.
1. What RustFS Is Today
| Aspect | Current State |
|---|---|
| Storage model | S3-compatible object storage (PUT/GET/DELETE whole objects) |
| Data path | HTTP → erasure coding → local disk I/O via ecstore (87K lines) |
| Protocol support | S3 API, OpenStack Swift, FTP/FTPS, WebDAV |
| Data durability | Erasure coding with bitrot protection and healing |
| Clustering | Distributed mode via gRPC inter-node RPC (still under testing) |
| Language | 98% Rust, Apache 2.0 license |
| Maturity | Beta (1.0.0-beta.7), GA planned July 2026 |
| POSIX support | None |
| Block device support | None |
| RDMA/DPU | Announced on roadmap, separate gpu-cache repo exists (early stage) |
2. What BeeGFS / Lustre Provide (The Target)
2.1 POSIX Parallel Filesystem Capabilities
AI training workloads (PyTorch DataLoader, TensorFlow tf.data, HuggingFace datasets) rely on:
fd = open("/mnt/cluster/dataset/shard-00042.bin", O_RDONLY);
fseek(fd, offset, SEEK_SET); // random seek into middle of file
fread(buf, chunk_size, 1, fd); // read a specific region
// ... repeat from many processes simultaneously
What BeeGFS/Lustre deliver that RustFS does not:
- POSIX VFS integration: mount as a kernel filesystem or FUSE, exposing
open/read/write/seek/close/stat/readdir/mmap - Byte-range I/O: read/write arbitrary byte ranges without downloading whole objects
- File striping: a single file is split across N storage targets; parallel I/O from all targets simultaneously
- Client-side caching / read-ahead: predictive prefetch for sequential and strided access patterns
- Distributed metadata: separate MDS (Metadata Server) from OSS (Object Storage Server)
- File locking:
flock()/fcntl()advisory and mandatory locks - Concurrent readers: hundreds of GPU nodes reading different offsets of the same file simultaneously with near-linear aggregate bandwidth scaling
2.2 Block Device Capabilities (VM / Bare-Metal Disks)
For VMs (KVM/QEMU, VMware, Hyper-V) and bare-metal iSCSI/NVMe-oF boot:
- Block-level I/O: fixed-size block reads/writes (typically 4K-64K), not object-level
- Thin provisioning: allocate-on-write with overcommit
- Snapshots & clones: instant copy-on-write snapshots
- Consistency: write-ordering guarantees, flush/barrier support
- Export protocols: iSCSI, NVMe-oF (NVMe over Fabrics), virtio-blk, NBD
- Examples: Ceph RBD, Longhorn, OpenEBS, LVM on shared storage
3. Gap Analysis: RustFS vs. Requirements
3.1 POSIX Filesystem Layer — MISSING ENTIRELY
| Requirement | RustFS Status | Gap Severity |
|---|---|---|
| Kernel VFS / FUSE mount point | Not implemented | CRITICAL |
open/read/write/seek/close syscalls | Objects are HTTP GET/PUT only | CRITICAL |
Byte-range read (fseek + fread) | S3 Range GET exists but no POSIX mapping | HIGH |
| Byte-range write (random write into file) | Not supported (objects are immutable-on-PUT) | CRITICAL |
| File striping across nodes | Erasure coding != striping for parallel read throughput | HIGH |
| Distributed metadata server | No separate MDS; metadata is in ecstore per-bucket | HIGH |
readdir / stat / chmod / POSIX permissions | Not implemented | CRITICAL |
| Client-side read-ahead / caching | No client component exists | HIGH |
mmap() support | Not possible without VFS layer | CRITICAL |
File locking (flock / fcntl) | Not implemented | MEDIUM |
3.2 Block Device Layer — MISSING ENTIRELY
| Requirement | RustFS Status | Gap Severity |
|---|---|---|
| Fixed-block I/O (4K aligned R/W) | Object-level only | CRITICAL |
| iSCSI / NVMe-oF target | Not implemented | CRITICAL |
| Thin provisioning | Not applicable to object model | CRITICAL |
| Snapshots & clones (COW) | S3 versioning exists, but no block-level COW | HIGH |
| Write barriers / flush semantics | HTTP has no equivalent | CRITICAL |
| NBD (Network Block Device) server | Not implemented | CRITICAL |
3.3 What RustFS Has That Can Be Reused
| Component | Reusable For |
|---|---|
ecstore erasure coding engine | Data durability for both POSIX and block layers |
io-core zero-copy I/O + buffer pool | Foundational I/O primitives |
rio reader pipeline (encrypt → compress → hash) | Data pipeline for all storage paths |
gRPC inter-node RPC (protos/) | Cluster communication backbone |
| IAM / policy engine | Multi-tenant access control |
lock distributed lock manager | Foundation for file locking |
| Observability stack (Prometheus, OTel) | Monitoring for new subsystems |
heal + scanner | Background integrity checking |
4. Proposed Architecture to Close the Gaps
4.1 Overall System Design
┌──────────────────────────────────────────────────────────┐
│ CLIENT NODES │
│ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ │
│ │ FUSE Client │ │ Kernel Client│ │ NBD / iSCSI │ │
│ │ (rustfs-fuse)│ │ (rustfs.ko) │ │ Client │ │
│ └──────┬───────┘ └──────┬───────┘ └──────┬───────┘ │
│ │ POSIX syscalls │ VFS ops │ Block I/O │
└─────────┼──────────────────┼─────────────────┼───────────┘
│ │ │
▼ ▼ ▼
┌──────────────────────────────────────────────────────────┐
│ NETWORK LAYER │
│ gRPC / RDMA (future) / TCP │
└────────────┬───────────────────────┬─────────────────────┘
│ │
▼ ▼
┌────────────────────┐ ┌─────────────────────┐
│ METADATA SERVER │ │ STORAGE SERVERS │
│ (rustfs-mds) │ │ (rustfs-oss) │
│ │ │ │
│ • Inode table │ │ • Chunk storage │
│ • Directory tree │ │ • Erasure coding │
│ • File→chunk map │ │ • Block volumes │
│ • POSIX attrs │ │ • Replication │
│ • Lock manager │ │ │
│ • Lease manager │ │ (reuses ecstore, │
│ │ │ io-core, rio) │
└────────────────────┘ └─────────────────────┘
4.2 POSIX Layer: What to Build
A. FUSE Client (rustfs-fuse) — Start Here
This is the fastest path to a working POSIX mount:
New crate: crates/fuse-client/
Dependencies: fuser (Rust FUSE library), gRPC client
Must implement:
├── init() → connect to MDS + discover OSS nodes
├── lookup() → MDS inode lookup
├── getattr() → MDS stat
├── readdir() → MDS directory listing
├── open() → MDS: allocate file handle, get chunk map
├── read() → OSS: parallel byte-range reads from striped chunks
├── write() → OSS: write to chunk targets, update MDS
├── seek() → client-local offset tracking
├── flush()/fsync() → OSS: ensure durability
├── release() → MDS: release file handle + leases
├── mkdir/rmdir/unlink/rename → MDS operations
└── setattr/chmod/chown → MDS POSIX attribute updates
Key design decisions for AI workloads:
- Stripe size: 1MB–4MB chunks (matches typical AI data loader read sizes)
- Read-ahead: adaptive prefetch based on sequential detection (like Lustre's
lliteread-ahead) - Client-side page cache: honor Linux VFS page cache for repeated reads
- Parallel chunk fetch: for a single
read(fd, buf, 64MB), fetch all stripe chunks in parallel from different OSS nodes
B. Metadata Server (rustfs-mds)
New crate: crates/mds/
Core data structures:
├── InodeTable → inode_id → {type, size, uid, gid, mode, timestamps, xattrs}
├── DirectoryTree → parent_inode + name → child_inode
├── ChunkMap → inode_id + offset → [(oss_node, chunk_id, stripe_idx)]
├── LockTable → inode_id → [lock_range, lock_type, owner]
├── LeaseManager → inode_id → [client_id, lease_type, expiry]
└── JournalLog → WAL for crash recovery
Storage backend options:
• RocksDB/sled for embedded single-MDS
• Raft consensus (via openraft) for HA MDS cluster
Metadata scalability (following Lustre's DNE model): partition the namespace into subtrees, each managed by a different MDS instance. This allows metadata operations to scale horizontally.
C. Object Storage Server Adaptation (rustfs-oss)
The existing ecstore needs a new interface alongside S3:
New trait: ChunkStorageService (gRPC)
rpc WriteChunk(chunk_id, offset, data) → Result
rpc ReadChunk(chunk_id, offset, length) → Stream<bytes>
rpc DeleteChunks(chunk_ids) → Result
rpc ChunkStat(chunk_id) → {size, checksum, location}
This reuses the existing ecstore erasure engine, rio pipeline, and io-core zero-copy I/O — but exposes chunk-level (not object-level) operations.
4.3 Block Device Layer: What to Build
A. Volume Manager (rustfs-volume)
New crate: crates/volume/
Manages block volumes as sequences of fixed-size extents:
Volume:
├── volume_id: UUID
├── size: u64 (logical size)
├── block_size: 4096
├── extent_size: 4MB
├── extent_map: BTreeMap<extent_offset, (oss_node, chunk_id)>
├── snapshot_tree: COW B-tree
└── provisioning: Thin | Thick
Operations:
├── create_volume(size, thin?) → volume_id
├── read_block(volume_id, lba, count) → data
├── write_block(volume_id, lba, data) → Result
├── flush(volume_id) → Result (write barriers)
├── snapshot(volume_id) → snapshot_id
├── clone(snapshot_id) → new_volume_id
└── delete_volume(volume_id) → Result
Thin provisioning: allocate extents on first write. The extent map starts empty; writes allocate new chunks on OSS nodes.
Copy-on-write snapshots: when creating a snapshot, freeze the extent map. Subsequent writes to the parent COW the modified extents to new locations.
B. Block Export Targets
iSCSI Target (widest compatibility — works with VMware, KVM, Windows, bare-metal):
New crate: crates/iscsi-target/
Dependency: existing Rust iSCSI target libraries or custom implementation
Maps:
SCSI READ(10/16) → volume.read_block()
SCSI WRITE(10/16) → volume.write_block()
SYNCHRONIZE_CACHE → volume.flush()
INQUIRY/READ_CAPACITY → volume metadata
NVMe-oF Target (highest performance — for RDMA-capable environments):
New crate: crates/nvmeof-target/
Requires: SPDK bindings or kernel NVMe target integration
Maps NVMe commands to volume operations.
Best combined with RDMA for lowest latency.
NBD Server (simplest — useful for Linux VMs and testing):
New crate: crates/nbd-server/
Simplest protocol: NBD_CMD_READ/WRITE/FLUSH/TRIM → volume ops
4.4 RDMA Integration (Future, Aligns with RustFS Roadmap)
RustFS already plans RDMA/DPU support. For the POSIX and block layers:
Replace gRPC chunk transfer with:
├── RDMA READ → zero-copy chunk fetch from OSS memory to client memory
├── RDMA WRITE → zero-copy chunk write from client to OSS
└── Benefits: bypass kernel network stack, ~1µs latency, 200+ Gbps
For the block layer specifically, RDMA + NVMe-oF gives near-local-disk latency, which is essential for VM boot disks.
5. AI Workload Compatibility: The "One Write, Multi fseek/fread" Pattern
This is the dominant pattern in AI training:
# Training data preparation (WRITE ONCE)
with open("/mnt/rustfs/dataset.bin", "wb") as f:
for shard in shards:
f.write(shard) # sequential write, striped across OSS nodes
# Training (MULTI-PROCESS RANDOM READ)
# 8 GPUs × 4 DataLoader workers = 32 concurrent readers
def worker(rank):
fd = open("/mnt/rustfs/dataset.bin", "rb")
for batch_idx in assigned_batches(rank):
fd.seek(batch_idx * batch_bytes) # random seek
data = fd.read(batch_bytes) # read batch
yield process(data)
Required optimizations for this pattern:
- Large stripe size (4MB): matches typical batch read sizes, avoids excessive metadata lookups
- Parallel stripe fetch: a single
read()spanning multiple stripes fetches from all OSS nodes simultaneously - Client read-ahead: detect sequential-within-stride patterns and prefetch next chunks
- Shared read leases: MDS grants shared-read leases to all clients on read-only files (no lock contention)
- Page cache cooperation: let Linux VFS cache recently-read chunks; critical for epoch-over-epoch data reuse
- Direct I/O option: for datasets that exceed RAM, bypass page cache and read straight from OSS via O_DIRECT
6. Implementation Roadmap
Phase 1: POSIX Foundation (3–4 months)
- Design and implement MDS with inode table, directory tree, chunk map
-
Build
rustfs-fuseclient with basicopen/read/write/seek/close/readdir/stat - Add chunk-level gRPC service to existing storage servers
- File striping: write stripes across N storage targets
-
Basic read: parallel chunk fetch for single
read()call -
Milestone: mount filesystem,
cpa file,catit back
Phase 2: AI-Ready POSIX (2–3 months)
- Client-side read-ahead and sequential detection
- Shared read leases for concurrent readers
- Large-file optimizations (adaptive stripe size, prefetch tuning)
-
O_DIRECTsupport for GPU-direct data loading - PyTorch DataLoader integration testing
- Milestone: train a model on 100+ GPU nodes reading from RustFS
Phase 3: Block Storage (3–4 months)
- Volume manager with extent mapping and thin provisioning
- COW snapshot engine
- NBD server (simplest export, for initial testing)
- iSCSI target (production VMs)
- Milestone: boot a VM from a RustFS block volume
Phase 4: Performance & Production (2–3 months)
- RDMA data path for chunk transfer (client ↔ OSS)
- NVMe-oF target for high-performance block export
- Kernel FUSE bypass (io_uring FUSE or native kernel module) for lower latency
- HA MDS with Raft consensus
- Distributed namespace (DNE-style metadata partitioning)
- Milestone: match BeeGFS throughput on standard benchmarks (IOR, mdtest)
Phase 5: Enterprise Features (ongoing)
- Unified namespace (S3 + POSIX + Block see same data where applicable)
- Quota management per-user / per-directory
- HSM (Hierarchical Storage Management): auto-tier cold data
- GPUDirect Storage integration (bypass CPU entirely)
- Kubernetes CSI driver (ReadWriteMany for POSIX, ReadWriteOnce for block)
7. Comparable Projects for Reference
| Project | Approach | Lesson for RustFS |
|---|---|---|
| Lustre | Kernel client + MDS + OSS, file striping | Gold standard for HPC POSIX; study client read-ahead design |
| BeeGFS | Userspace client (FUSE optional), MDS + OSS | Simpler than Lustre; good reference for metadata design |
| CephFS | FUSE + kernel client on top of RADOS object store | Proves object store → POSIX is viable; study MDS design |
| Ceph RBD | Block device on top of RADOS | Proves object store → block is viable; study COW snapshots |
| JuiceFS | FUSE client + any object store backend + Redis/TiKV MDS | Closest model — uses S3 as backend, adds POSIX via FUSE |
| SeaweedFS | Object store + FUSE mount + volume server | Shows the object-to-POSIX bridge pattern |
JuiceFS is the most directly relevant reference — it adds a POSIX layer on top of S3-compatible storage with a separate metadata engine. RustFS could follow a similar architecture but with tighter integration since it owns the storage engine.
8. Verdict
RustFS has strong foundational components (erasure coding, zero-copy I/O, Rust safety, gRPC cluster communication) that can be reused. But it is fundamentally an object store, and the two target use cases — POSIX parallel filesystem and block storage — require building multiple new subsystems that don't exist today:
| What's Needed | Estimated New Code | Difficulty |
|---|---|---|
| Metadata server (MDS) | ~15-25K lines | Hard |
| FUSE client | ~10-15K lines | Medium-Hard |
| Chunk-level gRPC service | ~3-5K lines | Medium |
| Volume manager + COW | ~10-15K lines | Hard |
| iSCSI/NVMe-oF/NBD targets | ~8-12K lines each | Medium-Hard |
| RDMA transport | ~5-10K lines | Hard |
Total new code: roughly 50-80K lines of Rust, comparable in size to the existing ecstore crate alone. This is a 12-18 month engineering effort for a team of 4-6 experienced systems programmers.
The alternative — and possibly faster path — is to use JuiceFS or a similar POSIX gateway in front of RustFS's S3 API for the POSIX use case, and integrate an existing block-storage layer (like Longhorn or a custom thin-provisioning layer) for block devices. This sacrifices tight integration but gets to production much faster.