RustFS 文档 文档

RustFS Unified Storage — Random Read Optimization

RustFS Unified Storage — Random Read Optimization

Document:    RUSTFS-LLD-002
Status:      DRAFT
Version:     0.1.0
Companion:   RUSTFS-FEAT-001 (Feature Spec), RUSTFS-LLD-001 (LLD)
Scope:       Optimizing random-access reads (fseek/fread, pread) for AI and
             database-like workloads where prefetch and page cache are
             ineffective.

0. The Problem Statement

Sequential read optimization (read-ahead, large stripes, page cache reuse) is solved in RUSTFS-LLD-001. Random reads are a different animal and dominate many real AI workloads:

Workload                          Access pattern
────────────────────────────────  ──────────────────────────────────────────
Shuffled training (epoch)         random sample of records each step
WebDataset / TFRecord / Parquet   seek to record offset, read N bytes
Embedding tables / lookups        random row reads from a huge matrix
KV-cache / vector DB              point reads at arbitrary offsets
Graph neighbor sampling           scattered reads following edges
RAG retrieval                     random chunk reads from a corpus
Memory-mapped model weights       page-fault-driven random reads
Database on block volume          B-tree / index random 4-16 KiB reads

0.1 Why Random Reads Are Hard Here

Problem 1 — Read-ahead is useless or harmful.
  Truly random next-offset is unpredictable. Prefetching the "next" chunk
  wastes bandwidth and pollutes cache.

Problem 2 — Each read pays full latency, unamortized.
  Sequential reads amortize MDS lookup over many bytes. A random 4 KiB read
  that triggers a chunk-map lookup pays ~1 ms MDS RTT for 4 KiB → 250 MB/s
  ceiling regardless of disk speed.

Problem 3 — Read amplification from large chunks + EC.
  Chunk size 4 MiB, EC block 1 MiB. A 4 KiB random read, naively handled,
  reads/reconstructs a 1 MiB EC block = 256x amplification.
  Across the network that's 1 MiB transferred to serve 4 KiB.

Problem 4 — Page cache thrashes when working set > RAM.
  Datasets are 10-100x larger than client RAM. LRU page cache hit rate on
  uniform-random access over a set N× larger than cache ≈ cache_size / N.
  For 10 TB dataset, 256 GB cache → ~2.5% hit rate. Cache is nearly useless
  unless access is skewed.

Problem 5 — IOPS, not bandwidth, is the bottleneck.
  A 4 KiB random read workload needs HIGH IOPS. NVMe can do ~1M IOPS but
  only if the I/O path has deep queues and low per-op overhead. A blocking,
  one-op-at-a-time path caps at ~10K IOPS.

0.2 The Five Optimization Pillars

1. Eliminate metadata from the read critical path   → layout lease + direct addr
2. Read only what's needed (kill amplification)      → EC partial read, sub-chunk
3. Make the few cacheable bytes count                → frequency cache + small tier
4. Turn "random" into "prefetchable" when possible   → access-plan hinting
5. Maximize IOPS on the I/O path                      → io_uring high-QD + coalescing

1. Pillar 1 — Remove Metadata From the Critical Path

1.1 Full Chunk Map on Open + Layout Lease

For random-read files, the client fetches the entire chunk map at open() and holds a Layout lease. After that, every random read resolves its chunk location purely in client memory — zero MDS traffic per read.

// crate: rustfs-fuse
// file:  crates/fuse-client/src/random_read.rs

/// On open() of a file flagged for random access, prefetch the full chunk map.
///
/// Cost: one MDS RPC returning the full ChunkMap (for a 10 GiB file with
/// 4 MiB chunks = 2560 entries ≈ 80 KiB on the wire). Amortized over
/// thousands of random reads → negligible.
///
/// The Layout lease guarantees these locations stay valid; the MDS will
/// revoke (callback) if the file is truncated, re-striped, or extended.
pub async fn open_for_random(&self, ino: InodeId) -> Result<OpenFile> {
    let resp = self.mds.get_chunk_map(GetChunkMapReq {
        ino,
        offset: 0,
        length: 0,          // 0 = entire file
        want_layout_lease: true,
    }).await?;

    Ok(OpenFile {
        ino,
        chunk_map: Arc::new(resp.chunk_map),   // full map cached
        lease: resp.layout_lease,              // Layout lease held
        access_hint: AccessHint::Random,
        ..Default::default()
    })
}

1.2 Access Pattern Detection and Hint Propagation

// Detect random access and switch strategy automatically.

pub enum AccessHint {
    Unknown,
    Sequential,   // → read-ahead engine (LLD-001 §6.1.3)
    Random,       // → this document's path
    AccessPlan,   // → §4, app provided an explicit plan
}

impl ReadAheadEngine {
    /// When detector classifies access as Random:
    ///   - disable prefetch (it would waste bandwidth)
    ///   - ensure full chunk map is resident
    ///   - switch cache policy to frequency-based (TinyLFU) admission
    fn on_random_detected(&self, ino: InodeId, file: &OpenFile) {
        self.disable_prefetch(ino);
        self.cache.set_policy(ino, CachePolicy::FrequencyAdmission);
    }
}

Application-level hints (fastest path) via posix_fadvise:

// Application code (PyTorch DataLoader, parquet reader, etc.)
posix_fadvise(fd, 0, 0, POSIX_FADV_RANDOM);    // → AccessHint::Random
posix_fadvise(fd, 0, 0, POSIX_FADV_WILLNEED);  // → prefetch specific range

The FUSE client honors fadvise to set the hint immediately rather than waiting for the detector to observe enough reads.


2. Pillar 2 — Read Only What's Needed

2.1 EC Partial Read (No Full-Stripe Reconstruction)

This is the single biggest win against amplification.

Standard EC read (naive): to read ANY byte of an EC-coded chunk, read all
  data shards, reconstruct the full block, then slice out the wanted bytes.
  4 KiB read → 1 MiB block reconstruction → 256x amplification.

Optimized EC read (degraded-free fast path): when all DATA shards are healthy
  (the common case), the wanted bytes live in a known subset of data shards.
  Reconstruction is ONLY needed when a data shard is missing/corrupt.

  Reed-Solomon systematic encoding: the first K shards ARE the original data
  (parity shards are extra). So a byte at file-offset O within a chunk lives
  in data shard (O / shard_stripe_size). Read just that shard's relevant
  bytes — no parity, no reconstruction.
// crate: rustfs-chunk-service
// file:  crates/chunk-service/src/ec_partial_read.rs

/// Partial read from an EC-coded chunk.
///
/// Systematic RS code: shards [0..K) are the literal data, [K..K+M) are parity.
/// To read bytes [offset, offset+len) within the chunk, we compute which
/// data shards and which byte ranges within them are needed, and read ONLY
/// those — bypassing reconstruction entirely when those shards are healthy.
///
/// Layout within a chunk for K data shards, stripe unit S (e.g., 64 KiB):
///   byte b → shard = (b / S) % K, shard_offset = (b / (S*K))*S + (b % S)
///   (RAID-0-style striping of the chunk across the K data shards)
pub async fn ec_partial_read(
    &self,
    chunk: &ChunkLayout,         // K, M, stripe_unit, shard locations
    offset: u64,                 // byte offset within the chunk
    length: u64,
) -> Result<Bytes> {
    // 1. Compute which (data_shard, shard_range) pieces cover [offset, offset+len)
    let pieces = map_chunk_range_to_shards(
        offset, length, chunk.k, chunk.stripe_unit
    );

    // 2. Read only the needed shard ranges, in parallel, from data shards.
    //    No parity shards touched. No decode step.
    let mut reads = Vec::with_capacity(pieces.len());
    for piece in &pieces {
        let loc = &chunk.data_shards[piece.shard_idx];
        if !self.is_healthy(loc) {
            // A needed data shard is down → fall back to reconstruction
            return self.ec_reconstruct_read(chunk, offset, length).await;
        }
        reads.push(self.disk_engine.read_chunk(
            loc.disk_id, loc.fd, piece.shard_offset, piece.len, /*direct*/ true,
        ));
    }
    let parts = futures::future::try_join_all(reads).await?;

    // 3. Reassemble in offset order (no XOR, no GF math — pure copy)
    Ok(reassemble_stripe(parts, &pieces))
}

/// Amplification analysis:
///   4 KiB read, stripe_unit 64 KiB:
///     - touches 1 data shard, reads 4 KiB from it
///     - amplification: 1x (vs 256x naive)
///   1 MiB read spanning K=4 shards, stripe 64 KiB:
///     - touches 4 shards, ~256 KiB each, 4 parallel reads
///     - amplification: 1x, full parallelism

Decision: systematic RS with chunk-internal striping at a tunable stripe_unit (default 64 KiB). Partial reads touch only the data shards covering the range. Reconstruction is the exception path (only when a needed data shard is unhealthy), not the default. This makes EC viable for random reads, not just sequential.

Tradeoff: smaller stripe_unit = less amplification for tiny reads but more shards touched for medium reads (more parallel I/O ops). 64 KiB balances both. For pure point-read workloads, set stripe_unit to the typical record size.

2.2 Sub-Chunk Read Granularity Everywhere

The ChunkService.ReadChunk RPC already takes (offset, length). The full path
honors byte-range granularity end to end:

  Application read(fd, buf, 4096) at offset 5_000_000
    → FUSE read(ino, off=5M, size=4096)
    → chunk_map.resolve_range(5M, 4096)
        → chunk_index = 5M / 4MiB = 1, offset_in_chunk = 5M - 4MiB = 808_064
    → ChunkService.ReadChunk(chunk_id, offset=808064, length=4096)
    → ec_partial_read reads 4 KiB from 1 data shard
    → 4 KiB returned

No point in this path reads more than 4 KiB (plus shard alignment rounding).

2.3 Alignment-Aware Reads (O_DIRECT)

O_DIRECT reads must be block-aligned (512 or 4096). For a misaligned random
read [5_000_000, 5_000_000+4096):
  - Round down start to 4096 boundary: 4_999_168 wait, compute properly:
    aligned_start = 5_000_000 & !4095 = 4_999_168? No: 5_000_000 / 4096 = 1220.7
    aligned_start = 1220 * 4096 = 4_997_120
    aligned_end   = ceil((5_000_000+4096)/4096)*4096 = 5_005_312? compute:
    (5_004_096)/4096 = 1221.99 → 1222*4096 = 5_005_312
  - Read aligned superset, slice out exact bytes.
  - Read amplification here is at most 2 extra blocks (8 KiB) — negligible.

The DiskIoEngine handles alignment transparently; callers pass exact ranges.

3. Pillar 3 — Make Cacheable Bytes Count

3.1 Sub-Chunk (Fragment) Cache

The client caches the bytes actually read, not whole chunks. Caching whole 4 MiB chunks for 4 KiB reads wastes 99.9% of cache capacity.

// crate: rustfs-fuse
// file:  crates/fuse-client/src/fragment_cache.rs

/// Fragment cache: caches arbitrary byte ranges, not fixed chunks.
///
/// Keyed by (inode, aligned_offset). Stores small aligned fragments
/// (default 64 KiB grains — matches EC stripe unit, balances metadata
/// overhead vs waste).
///
/// Eviction: TinyLFU (moka) — frequency-aware admission. A one-shot random
/// read does NOT evict a frequently-read fragment. This is critical for
/// skewed access (hot rows in an embedding table, popular records).
pub struct FragmentCache {
    grain_size: u64,                          // 64 KiB
    cache:      moka::sync::Cache<FragKey, Bytes>,
}

#[derive(Hash, Eq, PartialEq, Clone)]
struct FragKey {
    ino:    InodeId,
    grain:  u64,        // aligned_offset / grain_size
}

impl FragmentCache {
    /// Resolve a read against the fragment cache.
    /// Returns hits and the missing grain ranges to fetch.
    pub fn lookup(&self, ino: InodeId, offset: u64, len: u64) -> FragLookup {
        let first = offset / self.grain_size;
        let last  = (offset + len - 1) / self.grain_size;
        let mut hits = Vec::new();
        let mut misses = Vec::new();
        for g in first..=last {
            match self.cache.get(&FragKey { ino, grain: g }) {
                Some(data) => hits.push((g, data)),
                None => misses.push(g),
            }
        }
        FragLookup { hits, misses: coalesce(misses) }
    }

    pub fn insert(&self, ino: InodeId, grain: u64, data: Bytes) {
        self.cache.insert(FragKey { ino, grain }, data);
    }
}

Why TinyLFU matters for random reads: under skewed access (Zipfian, which most real workloads are), a frequency-admission cache massively outperforms LRU. LRU lets a flood of one-shot reads evict hot data ("cache scan pollution"); TinyLFU only admits an item if it's more frequent than the victim it would evict. For an embedding table where 5% of rows get 80% of accesses, TinyLFU keeps those hot rows resident.

3.2 Dedicated Small-Object / Hot-Fragment Memory Tier

For point-read-heavy workloads, add an optional in-memory tier on DSS nodes:

  DSS RAM tier (optional, sized by config):
    - LRU/TinyLFU cache of hot fragments in DSS memory
    - Served without touching disk → ~5 µs response
    - Populated on read; warmed by access-plan hints (§4)

  Read path with tiers:
    1. Client fragment cache (client RAM)       ~1 µs
    2. DSS hot-fragment tier (DSS RAM)          ~5 µs  + network
    3. DSS NVMe (io_uring, high QD)             ~80 µs + network
    4. EC reconstruction (only if shard down)   ~200 µs + network

3.3 Cache Coherency Under Random Writes

For read-only datasets (the AI training common case): clients hold ReadData +
Layout leases, fragment cache never goes stale, max hit rate.

For read-write random access (database on POSIX): a write to a fragment
revokes ReadData leases on that inode for other clients (or, with byte-range
leases — §5, only the affected range). Affected fragments are invalidated.

4. Pillar 4 — Turn "Random" Into "Prefetchable": Access-Plan Hinting

This is the highest-leverage optimization for AI training, because most "random" access in training is pseudo-random with a known plan.

4.1 The Insight

A shuffled training epoch is NOT truly random — it's a known permutation:

  indices = list(range(dataset_size))
  rng = Random(seed + epoch)
  rng.shuffle(indices)
  for idx in indices:           # ← the access plan is KNOWN in advance
      record = dataset[idx]     # random offset, but predictable sequence

If the application hands this permutation to the storage client BEFORE the
reads happen, the client can prefetch records in plan order — converting a
random workload into a (deeply pipelined) prefetchable one.

4.2 Access-Plan API

// crate: rustfs-client
// file:  crates/client/src/access_plan.rs

/// An access plan: an ordered list of (offset, length) reads the application
/// will perform. Submitted via an ioctl or an extended API before reading.
///
/// The client uses this to prefetch ahead of the application's read cursor,
/// hiding latency behind compute. Depth is tunable (how far ahead to prefetch).
#[derive(Clone, Debug)]
pub struct AccessPlan {
    pub ino:      InodeId,
    pub entries:  Vec<PlanEntry>,    // ordered reads
    pub prefetch_depth: usize,       // how many entries ahead to fetch (e.g., 64)
}

#[derive(Clone, Debug)]
pub struct PlanEntry {
    pub offset: u64,
    pub length: u64,
}

pub struct AccessPlanExecutor {
    plan:        AccessPlan,
    cursor:      AtomicUsize,        // app's current position in the plan
    prefetcher:  PrefetchPool,
    cache:       Arc<FragmentCache>,
}

impl AccessPlanExecutor {
    /// Called when the app reads. Advances the cursor and triggers prefetch
    /// of the next `prefetch_depth` entries.
    pub async fn on_read(&self, offset: u64, len: u64) -> Result<Bytes> {
        // Find this read in the plan, advance cursor
        let pos = self.cursor.fetch_add(1, Ordering::Relaxed);

        // Prefetch the look-ahead window in plan order, in parallel
        let window_end = (pos + self.plan.prefetch_depth).min(self.plan.entries.len());
        for i in (pos + 1)..window_end {
            let e = &self.plan.entries[i];
            if !self.cache.contains(self.plan.ino, e.offset, e.length) {
                self.prefetcher.submit(self.plan.ino, e.offset, e.length);
            }
        }

        // Serve the current read (likely already prefetched → cache hit)
        self.read_now(offset, len).await
    }
}

4.3 Framework Integration

# PyTorch integration via a custom Sampler + storage hint
# (provided as a small Python shim over the rustfs client)

class RustfsPlannedSampler(torch.utils.data.Sampler):
    def __init__(self, dataset, record_offsets, seed, epoch):
        self.offsets = record_offsets           # offset+len per record index
        self.order = list(range(len(dataset)))
        random.Random(seed + epoch).shuffle(self.order)
        # Hand the plan to the storage layer BEFORE iteration starts
        plan = [self.offsets[i] for i in self.order]
        rustfs_client.submit_access_plan(dataset.fd, plan, prefetch_depth=64)

    def __iter__(self):
        return iter(self.order)
Effect: each training step's record is prefetched ~64 steps ahead. With
typical step compute time (forward+backward) of a few ms and per-record fetch
latency < 1 ms, the storage latency is fully hidden. The "random" read becomes
effectively zero-latency from the app's perspective.

Without plan: each random read pays full fetch latency on the critical path.
With plan:    fetch latency overlaps compute; throughput becomes bandwidth-
              bound (not latency-bound).

4.4 Auto-Plan Detection (No App Changes)

When the app provides no explicit plan, the client can still detect structure:

  - Strided random (fixed-size records): detect record_size from read sizes,
    note that offsets are multiples of record_size → prefetch likely-next
    records based on observed stride histogram.
  - Index-file pattern: many formats (TFRecord index, parquet footer) read a
    small index region first, then seek to data offsets listed in it. The
    client can recognize the index read and pre-warm the listed offsets.

These are heuristic and lower-leverage than explicit plans, but require zero
application changes.

5. Pillar 5 — Maximize IOPS on the I/O Path

5.1 High Queue Depth io_uring for Random Reads

// crate: rustfs-chunk-service
// file:  crates/chunk-service/src/random_io.rs

/// Random read I/O path tuned for IOPS, not bandwidth.
///
/// Differences from the sequential path:
///   - Much higher queue depth (1024+ vs 128) to keep NVMe saturated
///     with many small in-flight ops.
///   - IORING_SETUP_IOPOLL for polled completions on NVMe (no interrupts,
///     lowest latency for small ops) — requires O_DIRECT.
///   - No read-ahead, no merging of non-adjacent reads.
///   - Batched submission: many SQEs submitted in one io_uring_enter.
pub struct RandomReadEngine {
    rings: Vec<PolledRing>,    // one per disk, IOPOLL mode
}

struct PolledRing {
    ring:        io_uring::IoUring,   // built with IORING_SETUP_IOPOLL
    queue_depth: u32,                 // 1024+
    poller:      JoinHandle<()>,      // dedicated completion poller task
}

impl RandomReadEngine {
    /// Submit many random reads at once; collect completions as they finish.
    pub async fn read_many(
        &self,
        reads: Vec<(u32 /*disk*/, RawFd, u64 /*off*/, u64 /*len*/)>,
    ) -> Result<Vec<Bytes>> {
        // Group by disk → submit a batch of SQEs per ring in one enter() call
        let by_disk = group_by_disk(reads);
        let mut futs = Vec::new();
        for (disk, ops) in by_disk {
            futs.push(self.rings[disk].submit_batch(ops));
        }
        let results = futures::future::try_join_all(futs).await?;
        Ok(flatten_in_order(results))
    }
}

Decision: dedicated polled (IOPOLL) io_uring rings for the random-read path with queue depth 1024+, O_DIRECT, batched SQE submission. This is what unlocks NVMe's ~1M IOPS. The sequential path keeps its buffered, lower-QD rings.

5.2 Read Coalescing and Scatter-Gather

When the access plan or concurrent app threads produce many small reads to the
SAME chunk that are ADJACENT or near-adjacent, coalesce them into one larger
read and scatter the result back:

  reads: [off=1000 len=512], [off=1600 len=512], [off=2200 len=512]
  → coalesce into one read [off=1000 len=1712] (gap-tolerant up to a threshold)
  → split the returned buffer back into the three requested ranges

This trades a little over-read (the gaps) for far fewer I/O ops. Gap threshold
default 16 KiB (coalesce if the wasted gap bytes < threshold).
// Coalescing logic
fn coalesce_reads(mut reads: Vec<ReadReq>, gap_threshold: u64) -> Vec<CoalescedRead> {
    reads.sort_by_key(|r| r.offset);
    let mut out = Vec::new();
    let mut cur: Option<CoalescedRead> = None;
    for r in reads {
        match &mut cur {
            Some(c) if r.offset <= c.end() + gap_threshold => {
                c.extend_to(r.offset + r.length);
                c.members.push(r);
            }
            _ => {
                if let Some(c) = cur.take() { out.push(c); }
                cur = Some(CoalescedRead::new(r));
            }
        }
    }
    if let Some(c) = cur { out.push(c); }
    out
}

5.3 RDMA Scatter-Gather for Batched Random Reads

RDMA supports scatter-gather lists (SGL): a single RDMA READ can gather
multiple non-contiguous remote regions into multiple local buffers, or a
single WR can carry an SGL.

For a batch of random reads to one DSS node, post a single RDMA operation
with an SGL describing all the wanted (remote_addr, len, local_buf) tuples.
One round trip serves N random reads.

  N random 4 KiB reads to DSS-3:
    Without SGL: N round trips (or N requests pipelined)
    With SGL:    1 RDMA op, N segments → 1 round trip

5.4 NUMA and Thread Affinity

High-IOPS random read paths are sensitive to cross-NUMA traffic:

  - Pin each disk's io_uring poller to a core on the NUMA node local to
    that NVMe's PCIe root complex.
  - Pin RDMA completion queue handlers to cores near the NIC.
  - Allocate fragment-cache and buffer-pool memory on the local NUMA node.
  - For DSS nodes that are also compute nodes, keep storage pollers off the
    cores reserved for the GPU feeding threads.

Config: dss.numa_pinning = "auto" | "manual" | "off"

6. End-to-End Random Read Paths

6.1 Best Case: Planned Read, Client Cache Hit

Application: fread(buf, 4096, 1, fp) at a planned offset
  │
  ▼ FUSE read(ino, off, 4096)  [FUSE-over-io_uring]                 ~1 µs
  │
  ▼ AccessPlanExecutor: this offset was prefetched 64 steps ago      0 µs
  │
  ▼ FragmentCache hit (client RAM)                                   ~1 µs
  │
  ▼ return 4 KiB
  ─────────────────────────────────────────────────────────────────
  TOTAL: ~2 µs  (no network, no MDS, no disk)

6.2 Common Case: Unplanned Random Read, Cache Miss, RDMA

Application: fread at a random offset
  │
  ▼ FUSE read(ino, off, 4096)                                        ~1 µs
  │
  ▼ Layout lease held → chunk location resolved in client RAM        ~0.5 µs
  │  (NO MDS round trip — this is the key win from Pillar 1)
  │
  ▼ FragmentCache miss → fetch grain (64 KiB) covering the read
  │
  ▼ pick replica (power-of-two)                                      ~0.2 µs
  │
  ▼ RDMA READ 64 KiB (or DSS hot-tier hit)                           ~8 µs
  │  └─ DSS: ec_partial_read → 1 data shard, 64 KiB, no reconstruct
  │
  ▼ cache grain, slice out 4 KiB, FUSE reply                         ~2 µs
  ─────────────────────────────────────────────────────────────────
  TOTAL: ~12 µs   (vs ~1.4 ms naive with MDS lookup + full-block EC)
                  → ~100x improvement on the cold-miss path

6.3 Worst Case: Cold, Degraded EC, TCP Fallback

  FUSE read                                                          ~5 µs
  Layout lease miss → MDS chunk map lookup                           ~1 ms
  EC reconstruction (one data shard down): read K shards + decode    ~300 µs
  io_uring TCP transfer                                              ~50 µs
  reply                                                              ~10 µs
  ─────────────────────────────────────────────────────────────────
  TOTAL: ~1.4 ms   (acceptable; this is the rare degraded path)

7. Configuration

# /etc/rustfs/random-read.toml

[random_read]
# Auto-fetch full chunk map + layout lease on open for files detected
# as random-access (or hinted via posix_fadvise(POSIX_FADV_RANDOM)).
prefetch_full_chunkmap = true
hold_layout_lease = true

[ec]
# Chunk-internal stripe unit. Smaller = less amplification for tiny reads.
# Set near your typical record size for pure point-read workloads.
stripe_unit = "64KiB"
# Use partial read (no reconstruction) when data shards are healthy.
partial_read_fast_path = true

[fragment_cache]
enabled = true
grain_size = "64KiB"          # cache granularity for random reads
capacity = "8GiB"             # per-client
policy = "tinylfu"            # frequency-aware admission

[dss_hot_tier]
enabled = true
capacity = "32GiB"            # per-DSS-node in-memory hot fragment tier
policy = "tinylfu"

[access_plan]
enabled = true
default_prefetch_depth = 64   # entries to prefetch ahead in plan order
auto_detect = true            # heuristic plan detection without app hints

[random_io]
queue_depth = 1024            # io_uring QD for random read rings
io_poll = true                # IOPOLL polled completions (needs O_DIRECT)
coalesce_gap_threshold = "16KiB"
numa_pinning = "auto"

8. New / Modified Components

Component                              New?  Crate                  Lines
─────────────────────────────────────  ────  ─────────────────────  ──────
EC partial read (no reconstruction)    NEW   chunk-service          ~2K
  + chunk-internal striping layout
RandomReadEngine (IOPOLL, high QD)     NEW   chunk-service          ~1.5K
FragmentCache (sub-chunk, TinyLFU)     NEW   fuse-client            ~1.5K
DSS hot-fragment memory tier           NEW   chunk-service          ~1K
AccessPlan API + executor              NEW   client                 ~2K
  + framework shims (Python/C)
Auto-plan / stride detector            NEW   client                 ~1K
Read coalescing + scatter-gather       NEW   client + chunk-service ~1K
RDMA SGL batched reads                 NEW   rdma                   ~1K
posix_fadvise handling                 MOD   fuse-client            +300
Layout-lease full-chunkmap-on-open     MOD   fuse-client            +400
NUMA pinning                           MOD   chunk-service          +500

9. Benchmarks and Acceptance Targets

Benchmark                              Target              Tool
─────────────────────────────────────  ──────────────────  ───────────────
4 KiB random read, cached (client)     > 2 M IOPS/client   fio randread
4 KiB random read, RDMA, EC partial    > 500 K IOPS/client fio randread
4 KiB random read latency p50 (RDMA)   < 15 µs             fio
4 KiB random read latency p99 (RDMA)   < 50 µs             fio (hedging caps tail)
EC partial read amplification (4 KiB)  < 1.2x              custom metric
Shuffled epoch w/ access plan vs       > 5x throughput     real training run
  no plan (latency hidden)
Embedding lookup (Zipfian, hot 5%)     > 80% cache hit     custom (TinyLFU)
mdtest-style metadata (open+stat)      > 200 K ops/s/client mdtest
Acceptance test — random read AI workload:
  $ fio --name=randread --rw=randread --bs=4k --iodepth=256 \
        --numjobs=16 --directory=/mnt/rustfs/buckets/dataset \
        --size=10G --direct=1 --runtime=60
  # Verify: > 500K IOPS aggregate, p99 < 50µs (RDMA), EC amplification < 1.2x

  $ python train_with_plan.py --data=/mnt/rustfs/.../imagenet --epochs=3
  # Verify: epoch time bandwidth-bound, not latency-bound
  # Verify: GPU utilization > 90% (storage not the bottleneck)

10. Summary: How Each Problem Is Solved

Problem (from §0.1)                     Solution
─────────────────────────────────────  ─────────────────────────────────────
1. Read-ahead useless for random        Access-plan hinting (§4): app-known
                                         permutation → prefetch in plan order;
                                         auto-detect for unhinted workloads

2. Per-read full latency unamortized     Layout lease + full chunk map on open
                                         (§1): zero MDS traffic per read →
                                         removes the 1 ms MDS RTT from path

3. Large-chunk + EC amplification         EC partial read with systematic RS +
                                         chunk-internal striping (§2.1):
                                         touch only needed data shards, no
                                         reconstruction → ~1x amplification

4. Page cache thrash (working set>RAM)    Sub-chunk fragment cache + TinyLFU
                                         frequency admission (§3): cache only
                                         the bytes read; keep hot fragments
                                         against scan pollution; DSS hot tier

5. IOPS bottleneck                        High-QD IOPOLL io_uring (§5.1),
                                         read coalescing (§5.2), RDMA SGL
                                         (§5.3), NUMA pinning (§5.4)

This document optimizes the random-read path specifically. It composes with RUSTFS-LLD-001 (which owns the sequential path and the general I/O stack): the access-pattern detector routes each open file to either the read-ahead engine (sequential) or the random-read path (this document), and a single file can switch modes if its access pattern changes.