RustFS Unified Storage — Random Read Optimization
RustFS Unified Storage — Random Read Optimization
Document: RUSTFS-LLD-002
Status: DRAFT
Version: 0.1.0
Companion: RUSTFS-FEAT-001 (Feature Spec), RUSTFS-LLD-001 (LLD)
Scope: Optimizing random-access reads (fseek/fread, pread) for AI and
database-like workloads where prefetch and page cache are
ineffective.
0. The Problem Statement
Sequential read optimization (read-ahead, large stripes, page cache reuse) is solved in RUSTFS-LLD-001. Random reads are a different animal and dominate many real AI workloads:
Workload Access pattern
──────────────────────────────── ──────────────────────────────────────────
Shuffled training (epoch) random sample of records each step
WebDataset / TFRecord / Parquet seek to record offset, read N bytes
Embedding tables / lookups random row reads from a huge matrix
KV-cache / vector DB point reads at arbitrary offsets
Graph neighbor sampling scattered reads following edges
RAG retrieval random chunk reads from a corpus
Memory-mapped model weights page-fault-driven random reads
Database on block volume B-tree / index random 4-16 KiB reads
0.1 Why Random Reads Are Hard Here
Problem 1 — Read-ahead is useless or harmful.
Truly random next-offset is unpredictable. Prefetching the "next" chunk
wastes bandwidth and pollutes cache.
Problem 2 — Each read pays full latency, unamortized.
Sequential reads amortize MDS lookup over many bytes. A random 4 KiB read
that triggers a chunk-map lookup pays ~1 ms MDS RTT for 4 KiB → 250 MB/s
ceiling regardless of disk speed.
Problem 3 — Read amplification from large chunks + EC.
Chunk size 4 MiB, EC block 1 MiB. A 4 KiB random read, naively handled,
reads/reconstructs a 1 MiB EC block = 256x amplification.
Across the network that's 1 MiB transferred to serve 4 KiB.
Problem 4 — Page cache thrashes when working set > RAM.
Datasets are 10-100x larger than client RAM. LRU page cache hit rate on
uniform-random access over a set N× larger than cache ≈ cache_size / N.
For 10 TB dataset, 256 GB cache → ~2.5% hit rate. Cache is nearly useless
unless access is skewed.
Problem 5 — IOPS, not bandwidth, is the bottleneck.
A 4 KiB random read workload needs HIGH IOPS. NVMe can do ~1M IOPS but
only if the I/O path has deep queues and low per-op overhead. A blocking,
one-op-at-a-time path caps at ~10K IOPS.
0.2 The Five Optimization Pillars
1. Eliminate metadata from the read critical path → layout lease + direct addr
2. Read only what's needed (kill amplification) → EC partial read, sub-chunk
3. Make the few cacheable bytes count → frequency cache + small tier
4. Turn "random" into "prefetchable" when possible → access-plan hinting
5. Maximize IOPS on the I/O path → io_uring high-QD + coalescing
1. Pillar 1 — Remove Metadata From the Critical Path
1.1 Full Chunk Map on Open + Layout Lease
For random-read files, the client fetches the entire chunk map at open() and holds a Layout lease. After that, every random read resolves its chunk location purely in client memory — zero MDS traffic per read.
// crate: rustfs-fuse
// file: crates/fuse-client/src/random_read.rs
/// On open() of a file flagged for random access, prefetch the full chunk map.
///
/// Cost: one MDS RPC returning the full ChunkMap (for a 10 GiB file with
/// 4 MiB chunks = 2560 entries ≈ 80 KiB on the wire). Amortized over
/// thousands of random reads → negligible.
///
/// The Layout lease guarantees these locations stay valid; the MDS will
/// revoke (callback) if the file is truncated, re-striped, or extended.
pub async fn open_for_random(&self, ino: InodeId) -> Result<OpenFile> {
let resp = self.mds.get_chunk_map(GetChunkMapReq {
ino,
offset: 0,
length: 0, // 0 = entire file
want_layout_lease: true,
}).await?;
Ok(OpenFile {
ino,
chunk_map: Arc::new(resp.chunk_map), // full map cached
lease: resp.layout_lease, // Layout lease held
access_hint: AccessHint::Random,
..Default::default()
})
}
1.2 Access Pattern Detection and Hint Propagation
// Detect random access and switch strategy automatically.
pub enum AccessHint {
Unknown,
Sequential, // → read-ahead engine (LLD-001 §6.1.3)
Random, // → this document's path
AccessPlan, // → §4, app provided an explicit plan
}
impl ReadAheadEngine {
/// When detector classifies access as Random:
/// - disable prefetch (it would waste bandwidth)
/// - ensure full chunk map is resident
/// - switch cache policy to frequency-based (TinyLFU) admission
fn on_random_detected(&self, ino: InodeId, file: &OpenFile) {
self.disable_prefetch(ino);
self.cache.set_policy(ino, CachePolicy::FrequencyAdmission);
}
}
Application-level hints (fastest path) via posix_fadvise:
// Application code (PyTorch DataLoader, parquet reader, etc.)
posix_fadvise(fd, 0, 0, POSIX_FADV_RANDOM); // → AccessHint::Random
posix_fadvise(fd, 0, 0, POSIX_FADV_WILLNEED); // → prefetch specific range
The FUSE client honors fadvise to set the hint immediately rather than waiting
for the detector to observe enough reads.
2. Pillar 2 — Read Only What's Needed
2.1 EC Partial Read (No Full-Stripe Reconstruction)
This is the single biggest win against amplification.
Standard EC read (naive): to read ANY byte of an EC-coded chunk, read all
data shards, reconstruct the full block, then slice out the wanted bytes.
4 KiB read → 1 MiB block reconstruction → 256x amplification.
Optimized EC read (degraded-free fast path): when all DATA shards are healthy
(the common case), the wanted bytes live in a known subset of data shards.
Reconstruction is ONLY needed when a data shard is missing/corrupt.
Reed-Solomon systematic encoding: the first K shards ARE the original data
(parity shards are extra). So a byte at file-offset O within a chunk lives
in data shard (O / shard_stripe_size). Read just that shard's relevant
bytes — no parity, no reconstruction.
// crate: rustfs-chunk-service
// file: crates/chunk-service/src/ec_partial_read.rs
/// Partial read from an EC-coded chunk.
///
/// Systematic RS code: shards [0..K) are the literal data, [K..K+M) are parity.
/// To read bytes [offset, offset+len) within the chunk, we compute which
/// data shards and which byte ranges within them are needed, and read ONLY
/// those — bypassing reconstruction entirely when those shards are healthy.
///
/// Layout within a chunk for K data shards, stripe unit S (e.g., 64 KiB):
/// byte b → shard = (b / S) % K, shard_offset = (b / (S*K))*S + (b % S)
/// (RAID-0-style striping of the chunk across the K data shards)
pub async fn ec_partial_read(
&self,
chunk: &ChunkLayout, // K, M, stripe_unit, shard locations
offset: u64, // byte offset within the chunk
length: u64,
) -> Result<Bytes> {
// 1. Compute which (data_shard, shard_range) pieces cover [offset, offset+len)
let pieces = map_chunk_range_to_shards(
offset, length, chunk.k, chunk.stripe_unit
);
// 2. Read only the needed shard ranges, in parallel, from data shards.
// No parity shards touched. No decode step.
let mut reads = Vec::with_capacity(pieces.len());
for piece in &pieces {
let loc = &chunk.data_shards[piece.shard_idx];
if !self.is_healthy(loc) {
// A needed data shard is down → fall back to reconstruction
return self.ec_reconstruct_read(chunk, offset, length).await;
}
reads.push(self.disk_engine.read_chunk(
loc.disk_id, loc.fd, piece.shard_offset, piece.len, /*direct*/ true,
));
}
let parts = futures::future::try_join_all(reads).await?;
// 3. Reassemble in offset order (no XOR, no GF math — pure copy)
Ok(reassemble_stripe(parts, &pieces))
}
/// Amplification analysis:
/// 4 KiB read, stripe_unit 64 KiB:
/// - touches 1 data shard, reads 4 KiB from it
/// - amplification: 1x (vs 256x naive)
/// 1 MiB read spanning K=4 shards, stripe 64 KiB:
/// - touches 4 shards, ~256 KiB each, 4 parallel reads
/// - amplification: 1x, full parallelism
Decision: systematic RS with chunk-internal striping at a tunable
stripe_unit (default 64 KiB). Partial reads touch only the data shards
covering the range. Reconstruction is the exception path (only when a needed
data shard is unhealthy), not the default. This makes EC viable for random
reads, not just sequential.
Tradeoff: smaller stripe_unit = less amplification for tiny reads but more
shards touched for medium reads (more parallel I/O ops). 64 KiB balances both.
For pure point-read workloads, set stripe_unit to the typical record size.
2.2 Sub-Chunk Read Granularity Everywhere
The ChunkService.ReadChunk RPC already takes (offset, length). The full path
honors byte-range granularity end to end:
Application read(fd, buf, 4096) at offset 5_000_000
→ FUSE read(ino, off=5M, size=4096)
→ chunk_map.resolve_range(5M, 4096)
→ chunk_index = 5M / 4MiB = 1, offset_in_chunk = 5M - 4MiB = 808_064
→ ChunkService.ReadChunk(chunk_id, offset=808064, length=4096)
→ ec_partial_read reads 4 KiB from 1 data shard
→ 4 KiB returned
No point in this path reads more than 4 KiB (plus shard alignment rounding).
2.3 Alignment-Aware Reads (O_DIRECT)
O_DIRECT reads must be block-aligned (512 or 4096). For a misaligned random
read [5_000_000, 5_000_000+4096):
- Round down start to 4096 boundary: 4_999_168 wait, compute properly:
aligned_start = 5_000_000 & !4095 = 4_999_168? No: 5_000_000 / 4096 = 1220.7
aligned_start = 1220 * 4096 = 4_997_120
aligned_end = ceil((5_000_000+4096)/4096)*4096 = 5_005_312? compute:
(5_004_096)/4096 = 1221.99 → 1222*4096 = 5_005_312
- Read aligned superset, slice out exact bytes.
- Read amplification here is at most 2 extra blocks (8 KiB) — negligible.
The DiskIoEngine handles alignment transparently; callers pass exact ranges.
3. Pillar 3 — Make Cacheable Bytes Count
3.1 Sub-Chunk (Fragment) Cache
The client caches the bytes actually read, not whole chunks. Caching whole 4 MiB chunks for 4 KiB reads wastes 99.9% of cache capacity.
// crate: rustfs-fuse
// file: crates/fuse-client/src/fragment_cache.rs
/// Fragment cache: caches arbitrary byte ranges, not fixed chunks.
///
/// Keyed by (inode, aligned_offset). Stores small aligned fragments
/// (default 64 KiB grains — matches EC stripe unit, balances metadata
/// overhead vs waste).
///
/// Eviction: TinyLFU (moka) — frequency-aware admission. A one-shot random
/// read does NOT evict a frequently-read fragment. This is critical for
/// skewed access (hot rows in an embedding table, popular records).
pub struct FragmentCache {
grain_size: u64, // 64 KiB
cache: moka::sync::Cache<FragKey, Bytes>,
}
#[derive(Hash, Eq, PartialEq, Clone)]
struct FragKey {
ino: InodeId,
grain: u64, // aligned_offset / grain_size
}
impl FragmentCache {
/// Resolve a read against the fragment cache.
/// Returns hits and the missing grain ranges to fetch.
pub fn lookup(&self, ino: InodeId, offset: u64, len: u64) -> FragLookup {
let first = offset / self.grain_size;
let last = (offset + len - 1) / self.grain_size;
let mut hits = Vec::new();
let mut misses = Vec::new();
for g in first..=last {
match self.cache.get(&FragKey { ino, grain: g }) {
Some(data) => hits.push((g, data)),
None => misses.push(g),
}
}
FragLookup { hits, misses: coalesce(misses) }
}
pub fn insert(&self, ino: InodeId, grain: u64, data: Bytes) {
self.cache.insert(FragKey { ino, grain }, data);
}
}
Why TinyLFU matters for random reads: under skewed access (Zipfian, which most real workloads are), a frequency-admission cache massively outperforms LRU. LRU lets a flood of one-shot reads evict hot data ("cache scan pollution"); TinyLFU only admits an item if it's more frequent than the victim it would evict. For an embedding table where 5% of rows get 80% of accesses, TinyLFU keeps those hot rows resident.
3.2 Dedicated Small-Object / Hot-Fragment Memory Tier
For point-read-heavy workloads, add an optional in-memory tier on DSS nodes:
DSS RAM tier (optional, sized by config):
- LRU/TinyLFU cache of hot fragments in DSS memory
- Served without touching disk → ~5 µs response
- Populated on read; warmed by access-plan hints (§4)
Read path with tiers:
1. Client fragment cache (client RAM) ~1 µs
2. DSS hot-fragment tier (DSS RAM) ~5 µs + network
3. DSS NVMe (io_uring, high QD) ~80 µs + network
4. EC reconstruction (only if shard down) ~200 µs + network
3.3 Cache Coherency Under Random Writes
For read-only datasets (the AI training common case): clients hold ReadData +
Layout leases, fragment cache never goes stale, max hit rate.
For read-write random access (database on POSIX): a write to a fragment
revokes ReadData leases on that inode for other clients (or, with byte-range
leases — §5, only the affected range). Affected fragments are invalidated.
4. Pillar 4 — Turn "Random" Into "Prefetchable": Access-Plan Hinting
This is the highest-leverage optimization for AI training, because most "random" access in training is pseudo-random with a known plan.
4.1 The Insight
A shuffled training epoch is NOT truly random — it's a known permutation:
indices = list(range(dataset_size))
rng = Random(seed + epoch)
rng.shuffle(indices)
for idx in indices: # ← the access plan is KNOWN in advance
record = dataset[idx] # random offset, but predictable sequence
If the application hands this permutation to the storage client BEFORE the
reads happen, the client can prefetch records in plan order — converting a
random workload into a (deeply pipelined) prefetchable one.
4.2 Access-Plan API
// crate: rustfs-client
// file: crates/client/src/access_plan.rs
/// An access plan: an ordered list of (offset, length) reads the application
/// will perform. Submitted via an ioctl or an extended API before reading.
///
/// The client uses this to prefetch ahead of the application's read cursor,
/// hiding latency behind compute. Depth is tunable (how far ahead to prefetch).
#[derive(Clone, Debug)]
pub struct AccessPlan {
pub ino: InodeId,
pub entries: Vec<PlanEntry>, // ordered reads
pub prefetch_depth: usize, // how many entries ahead to fetch (e.g., 64)
}
#[derive(Clone, Debug)]
pub struct PlanEntry {
pub offset: u64,
pub length: u64,
}
pub struct AccessPlanExecutor {
plan: AccessPlan,
cursor: AtomicUsize, // app's current position in the plan
prefetcher: PrefetchPool,
cache: Arc<FragmentCache>,
}
impl AccessPlanExecutor {
/// Called when the app reads. Advances the cursor and triggers prefetch
/// of the next `prefetch_depth` entries.
pub async fn on_read(&self, offset: u64, len: u64) -> Result<Bytes> {
// Find this read in the plan, advance cursor
let pos = self.cursor.fetch_add(1, Ordering::Relaxed);
// Prefetch the look-ahead window in plan order, in parallel
let window_end = (pos + self.plan.prefetch_depth).min(self.plan.entries.len());
for i in (pos + 1)..window_end {
let e = &self.plan.entries[i];
if !self.cache.contains(self.plan.ino, e.offset, e.length) {
self.prefetcher.submit(self.plan.ino, e.offset, e.length);
}
}
// Serve the current read (likely already prefetched → cache hit)
self.read_now(offset, len).await
}
}
4.3 Framework Integration
# PyTorch integration via a custom Sampler + storage hint
# (provided as a small Python shim over the rustfs client)
class RustfsPlannedSampler(torch.utils.data.Sampler):
def __init__(self, dataset, record_offsets, seed, epoch):
self.offsets = record_offsets # offset+len per record index
self.order = list(range(len(dataset)))
random.Random(seed + epoch).shuffle(self.order)
# Hand the plan to the storage layer BEFORE iteration starts
plan = [self.offsets[i] for i in self.order]
rustfs_client.submit_access_plan(dataset.fd, plan, prefetch_depth=64)
def __iter__(self):
return iter(self.order)
Effect: each training step's record is prefetched ~64 steps ahead. With
typical step compute time (forward+backward) of a few ms and per-record fetch
latency < 1 ms, the storage latency is fully hidden. The "random" read becomes
effectively zero-latency from the app's perspective.
Without plan: each random read pays full fetch latency on the critical path.
With plan: fetch latency overlaps compute; throughput becomes bandwidth-
bound (not latency-bound).
4.4 Auto-Plan Detection (No App Changes)
When the app provides no explicit plan, the client can still detect structure:
- Strided random (fixed-size records): detect record_size from read sizes,
note that offsets are multiples of record_size → prefetch likely-next
records based on observed stride histogram.
- Index-file pattern: many formats (TFRecord index, parquet footer) read a
small index region first, then seek to data offsets listed in it. The
client can recognize the index read and pre-warm the listed offsets.
These are heuristic and lower-leverage than explicit plans, but require zero
application changes.
5. Pillar 5 — Maximize IOPS on the I/O Path
5.1 High Queue Depth io_uring for Random Reads
// crate: rustfs-chunk-service
// file: crates/chunk-service/src/random_io.rs
/// Random read I/O path tuned for IOPS, not bandwidth.
///
/// Differences from the sequential path:
/// - Much higher queue depth (1024+ vs 128) to keep NVMe saturated
/// with many small in-flight ops.
/// - IORING_SETUP_IOPOLL for polled completions on NVMe (no interrupts,
/// lowest latency for small ops) — requires O_DIRECT.
/// - No read-ahead, no merging of non-adjacent reads.
/// - Batched submission: many SQEs submitted in one io_uring_enter.
pub struct RandomReadEngine {
rings: Vec<PolledRing>, // one per disk, IOPOLL mode
}
struct PolledRing {
ring: io_uring::IoUring, // built with IORING_SETUP_IOPOLL
queue_depth: u32, // 1024+
poller: JoinHandle<()>, // dedicated completion poller task
}
impl RandomReadEngine {
/// Submit many random reads at once; collect completions as they finish.
pub async fn read_many(
&self,
reads: Vec<(u32 /*disk*/, RawFd, u64 /*off*/, u64 /*len*/)>,
) -> Result<Vec<Bytes>> {
// Group by disk → submit a batch of SQEs per ring in one enter() call
let by_disk = group_by_disk(reads);
let mut futs = Vec::new();
for (disk, ops) in by_disk {
futs.push(self.rings[disk].submit_batch(ops));
}
let results = futures::future::try_join_all(futs).await?;
Ok(flatten_in_order(results))
}
}
Decision: dedicated polled (IOPOLL) io_uring rings for the random-read path with queue depth 1024+, O_DIRECT, batched SQE submission. This is what unlocks NVMe's ~1M IOPS. The sequential path keeps its buffered, lower-QD rings.
5.2 Read Coalescing and Scatter-Gather
When the access plan or concurrent app threads produce many small reads to the
SAME chunk that are ADJACENT or near-adjacent, coalesce them into one larger
read and scatter the result back:
reads: [off=1000 len=512], [off=1600 len=512], [off=2200 len=512]
→ coalesce into one read [off=1000 len=1712] (gap-tolerant up to a threshold)
→ split the returned buffer back into the three requested ranges
This trades a little over-read (the gaps) for far fewer I/O ops. Gap threshold
default 16 KiB (coalesce if the wasted gap bytes < threshold).
// Coalescing logic
fn coalesce_reads(mut reads: Vec<ReadReq>, gap_threshold: u64) -> Vec<CoalescedRead> {
reads.sort_by_key(|r| r.offset);
let mut out = Vec::new();
let mut cur: Option<CoalescedRead> = None;
for r in reads {
match &mut cur {
Some(c) if r.offset <= c.end() + gap_threshold => {
c.extend_to(r.offset + r.length);
c.members.push(r);
}
_ => {
if let Some(c) = cur.take() { out.push(c); }
cur = Some(CoalescedRead::new(r));
}
}
}
if let Some(c) = cur { out.push(c); }
out
}
5.3 RDMA Scatter-Gather for Batched Random Reads
RDMA supports scatter-gather lists (SGL): a single RDMA READ can gather
multiple non-contiguous remote regions into multiple local buffers, or a
single WR can carry an SGL.
For a batch of random reads to one DSS node, post a single RDMA operation
with an SGL describing all the wanted (remote_addr, len, local_buf) tuples.
One round trip serves N random reads.
N random 4 KiB reads to DSS-3:
Without SGL: N round trips (or N requests pipelined)
With SGL: 1 RDMA op, N segments → 1 round trip
5.4 NUMA and Thread Affinity
High-IOPS random read paths are sensitive to cross-NUMA traffic:
- Pin each disk's io_uring poller to a core on the NUMA node local to
that NVMe's PCIe root complex.
- Pin RDMA completion queue handlers to cores near the NIC.
- Allocate fragment-cache and buffer-pool memory on the local NUMA node.
- For DSS nodes that are also compute nodes, keep storage pollers off the
cores reserved for the GPU feeding threads.
Config: dss.numa_pinning = "auto" | "manual" | "off"
6. End-to-End Random Read Paths
6.1 Best Case: Planned Read, Client Cache Hit
Application: fread(buf, 4096, 1, fp) at a planned offset
│
▼ FUSE read(ino, off, 4096) [FUSE-over-io_uring] ~1 µs
│
▼ AccessPlanExecutor: this offset was prefetched 64 steps ago 0 µs
│
▼ FragmentCache hit (client RAM) ~1 µs
│
▼ return 4 KiB
─────────────────────────────────────────────────────────────────
TOTAL: ~2 µs (no network, no MDS, no disk)
6.2 Common Case: Unplanned Random Read, Cache Miss, RDMA
Application: fread at a random offset
│
▼ FUSE read(ino, off, 4096) ~1 µs
│
▼ Layout lease held → chunk location resolved in client RAM ~0.5 µs
│ (NO MDS round trip — this is the key win from Pillar 1)
│
▼ FragmentCache miss → fetch grain (64 KiB) covering the read
│
▼ pick replica (power-of-two) ~0.2 µs
│
▼ RDMA READ 64 KiB (or DSS hot-tier hit) ~8 µs
│ └─ DSS: ec_partial_read → 1 data shard, 64 KiB, no reconstruct
│
▼ cache grain, slice out 4 KiB, FUSE reply ~2 µs
─────────────────────────────────────────────────────────────────
TOTAL: ~12 µs (vs ~1.4 ms naive with MDS lookup + full-block EC)
→ ~100x improvement on the cold-miss path
6.3 Worst Case: Cold, Degraded EC, TCP Fallback
FUSE read ~5 µs
Layout lease miss → MDS chunk map lookup ~1 ms
EC reconstruction (one data shard down): read K shards + decode ~300 µs
io_uring TCP transfer ~50 µs
reply ~10 µs
─────────────────────────────────────────────────────────────────
TOTAL: ~1.4 ms (acceptable; this is the rare degraded path)
7. Configuration
# /etc/rustfs/random-read.toml
[random_read]
# Auto-fetch full chunk map + layout lease on open for files detected
# as random-access (or hinted via posix_fadvise(POSIX_FADV_RANDOM)).
prefetch_full_chunkmap = true
hold_layout_lease = true
[ec]
# Chunk-internal stripe unit. Smaller = less amplification for tiny reads.
# Set near your typical record size for pure point-read workloads.
stripe_unit = "64KiB"
# Use partial read (no reconstruction) when data shards are healthy.
partial_read_fast_path = true
[fragment_cache]
enabled = true
grain_size = "64KiB" # cache granularity for random reads
capacity = "8GiB" # per-client
policy = "tinylfu" # frequency-aware admission
[dss_hot_tier]
enabled = true
capacity = "32GiB" # per-DSS-node in-memory hot fragment tier
policy = "tinylfu"
[access_plan]
enabled = true
default_prefetch_depth = 64 # entries to prefetch ahead in plan order
auto_detect = true # heuristic plan detection without app hints
[random_io]
queue_depth = 1024 # io_uring QD for random read rings
io_poll = true # IOPOLL polled completions (needs O_DIRECT)
coalesce_gap_threshold = "16KiB"
numa_pinning = "auto"
8. New / Modified Components
Component New? Crate Lines
───────────────────────────────────── ──── ───────────────────── ──────
EC partial read (no reconstruction) NEW chunk-service ~2K
+ chunk-internal striping layout
RandomReadEngine (IOPOLL, high QD) NEW chunk-service ~1.5K
FragmentCache (sub-chunk, TinyLFU) NEW fuse-client ~1.5K
DSS hot-fragment memory tier NEW chunk-service ~1K
AccessPlan API + executor NEW client ~2K
+ framework shims (Python/C)
Auto-plan / stride detector NEW client ~1K
Read coalescing + scatter-gather NEW client + chunk-service ~1K
RDMA SGL batched reads NEW rdma ~1K
posix_fadvise handling MOD fuse-client +300
Layout-lease full-chunkmap-on-open MOD fuse-client +400
NUMA pinning MOD chunk-service +500
9. Benchmarks and Acceptance Targets
Benchmark Target Tool
───────────────────────────────────── ────────────────── ───────────────
4 KiB random read, cached (client) > 2 M IOPS/client fio randread
4 KiB random read, RDMA, EC partial > 500 K IOPS/client fio randread
4 KiB random read latency p50 (RDMA) < 15 µs fio
4 KiB random read latency p99 (RDMA) < 50 µs fio (hedging caps tail)
EC partial read amplification (4 KiB) < 1.2x custom metric
Shuffled epoch w/ access plan vs > 5x throughput real training run
no plan (latency hidden)
Embedding lookup (Zipfian, hot 5%) > 80% cache hit custom (TinyLFU)
mdtest-style metadata (open+stat) > 200 K ops/s/client mdtest
Acceptance test — random read AI workload:
$ fio --name=randread --rw=randread --bs=4k --iodepth=256 \
--numjobs=16 --directory=/mnt/rustfs/buckets/dataset \
--size=10G --direct=1 --runtime=60
# Verify: > 500K IOPS aggregate, p99 < 50µs (RDMA), EC amplification < 1.2x
$ python train_with_plan.py --data=/mnt/rustfs/.../imagenet --epochs=3
# Verify: epoch time bandwidth-bound, not latency-bound
# Verify: GPU utilization > 90% (storage not the bottleneck)
10. Summary: How Each Problem Is Solved
Problem (from §0.1) Solution
───────────────────────────────────── ─────────────────────────────────────
1. Read-ahead useless for random Access-plan hinting (§4): app-known
permutation → prefetch in plan order;
auto-detect for unhinted workloads
2. Per-read full latency unamortized Layout lease + full chunk map on open
(§1): zero MDS traffic per read →
removes the 1 ms MDS RTT from path
3. Large-chunk + EC amplification EC partial read with systematic RS +
chunk-internal striping (§2.1):
touch only needed data shards, no
reconstruction → ~1x amplification
4. Page cache thrash (working set>RAM) Sub-chunk fragment cache + TinyLFU
frequency admission (§3): cache only
the bytes read; keep hot fragments
against scan pollution; DSS hot tier
5. IOPS bottleneck High-QD IOPOLL io_uring (§5.1),
read coalescing (§5.2), RDMA SGL
(§5.3), NUMA pinning (§5.4)
This document optimizes the random-read path specifically. It composes with RUSTFS-LLD-001 (which owns the sequential path and the general I/O stack): the access-pattern detector routes each open file to either the read-ahead engine (sequential) or the random-read path (this document), and a single file can switch modes if its access pattern changes.