Cloud Storage FUSE, the open source gcsfuse adapter, mounts a Cloud Storage bucket as a directory. Programs that expect files, such as PyTorch data loaders, model servers loading weights, or legacy batch tools, can then read and write objects with ordinary open and read calls. It is how GKE, Cloud Run and Vertex AI workloads commonly reach bucket data without copying it first.

The convenience hides a translation layer, and most problems with gcsfuse come from forgetting it is there. A bucket is not a disk: each file operation turns into one or more HTTP requests with object-store latency and per-operation cost, and several POSIX guarantees do not hold. This article explains how that translation works, how the caches and write paths change it, and how to configure it for the workloads it suits. Details come from Google's Cloud Storage FUSE documentation and the gcsfuse semantics guide as of October 2026. For Cloud Storage itself, see Google Cloud Storage.

Advertisement

How a file call becomes an object request

FUSE (Filesystem in Userspace) lets a normal process implement a file system. When an application calls read on a path under the mount, the kernel's VFS layer turns it into a request on /dev/fuse, the gcsfuse daemon picks it up, and the daemon decides how to answer it: from a cache, or with a Cloud Storage API call. The reply travels back through the kernel to the application, which sees a normal return value.

One read through Cloud Storage FUSE: kernel, daemon, caches, bucketApplicationopen, read, statLinux kernelVFS and page cacheFUSE device/dev/fuse requestsgcsfuse daemonuser space processMetadata cachesstat, type, negativeFile cachelocal SSD or RAM, LRUCloud StorageGET, list, uploadWrite bufferstreaming or stagedsyscalllookuphitmisscloseEvery cache miss is an HTTP request with object-store latency, never a local disk read.
The daemon sits between the kernel and the bucket. Caches decide whether a call costs a network round trip.
File operationWhat gcsfuse doesCost to watch
stat, openFetch object metadata, unless the stat cache holds itA request per uncached lookup
readRanged GET on the object, or a read from the file cacheLatency on every miss
readdir, lsObjects.list over the prefix, possibly several pagesSeveral list calls for big directories
write, closeStream or stage bytes, then create a new object generationUpload per closed file
Directory renameFast and atomic on hierarchical namespace buckets; in flat buckets allowed only up to rename-dir-limitOne operation per object in flat buckets

Each of these maps to Cloud Storage operations that are billed as Class A or Class B, so a tool that stats thousands of files in a loop costs money as well as time. The practical rule is that gcsfuse is fast for large sequential reads and writes, and slow and expensive for many small metadata operations.

The POSIX gaps

Google's documentation states plainly that Cloud Storage FUSE is not POSIX compliant. The gaps that matter in practice:

  • Directories are inferred. Cloud Storage has objects with slashes in their names. In a flat-namespace bucket, gcsfuse cannot see a directory that only exists implicitly, because objects were written under a/b/c.txt without an a/b/ placeholder, unless you enable implicit-dirs. Enabling it costs extra list calls on lookups.
  • No locking and no hard links. File locks and hard links are unsupported, so databases and version control repositories do not belong on a mount.
  • Concurrent writers. gcsfuse uses generation preconditions. When several mounts write the same object, the first to finish wins, and the others get ESTALE when they sync or close.
  • Directory renames. In flat buckets a directory rename is many object operations and not atomic, and it is refused beyond rename-dir-limit, which defaults to 0. Hierarchical namespace buckets make it fast and atomic.

Design around these rather than fighting them: one writer per object, immutable file names for published data, and a marker file written last to signal completeness.

Advertisement

Read path and caches

Without caches, every stat and every read miss is a round trip to Cloud Storage. gcsfuse has several caches, each trading consistency for speed.

The metadata caches remember object attributes (the stat cache), whether a name is a file or directory (the type cache), and names that do not exist (a negative cache). In the config file they sit under metadata-cache: ttl-secs defaults to 60 seconds, stat-cache-max-size-mb to 34, and negative-ttl-secs to 5. The semantics guide warns that stat caching breaks consistency unless the bucket is never modified, or only modified through a single mount. For immutable datasets, a TTL of -1 (never expire) removes almost all metadata traffic.

The list cache, file-system:kernel-list-cache-ttl-secs, lets the kernel keep directory listings; it defaults to 0, off. It helps data loaders that list the same large directory every epoch.

The file cache keeps whole object contents on local storage. It turns on when you set cache-dir (or on GKE, through the CSI driver). file-cache:max-size-mb of -1 means unlimited, bounded only by the disk, and 0 disables it; beyond the limit, entries are evicted least recently used first. By default only sequential reads from the start of a file populate the cache; cache-file-for-range-read makes a random first read fetch the whole object in the background, which suits formats like Parquet or safetensors that read a footer or header first. Cached content is still checked against the object generation once the metadata TTL expires. From version 2.12, parallel downloads are on by default when the file cache is enabled: several workers fetch chunks of one large file concurrently, controlled by parallel-downloads-per-file (default 16) and download-chunk-size-mb (default 200).

# /etc/gcsfuse/training.yaml: read-mostly dataset, one writer elsewhere
implicit-dirs: true              # the bucket was filled by gsutil or pipelines, not by gcsfuse
cache-dir: /mnt/localssd/gcsfuse # enables the file cache outside GKE
file-cache:
  max-size-mb: 1400000           # about 1.4 TB; -1 means "until the disk is full"
  cache-file-for-range-read: true
  enable-parallel-downloads: true
metadata-cache:
  ttl-secs: 3600                 # dataset is immutable during a run
  stat-cache-max-size-mb: 256
file-system:
  kernel-list-cache-ttl-secs: 3600
logging:
  severity: warning
mkdir -p /mnt/data
gcsfuse --config-file=/etc/gcsfuse/training.yaml --only-dir=datasets/v7 \
        acme-ml-data /mnt/data
# unmount
fusermount -u /mnt/data

Write path: streaming and staged

Since gcsfuse 3.0.0, streaming writes are the default (write:enable-streaming-writes). Sequential writes to a new file are uploaded to Cloud Storage as they happen, in memory buffers, instead of being staged in full on local disk first. That removes the local-disk requirement for large outputs and cuts the long pause at close that staged writes cause. Google's Cloud Run documentation puts the memory cost at about 64 MiB per file open for streaming writes, which adds up when a process writes many files at once.

Two rules follow from the semantics guide. First, fsync does not finalize the object under streaming writes; the object is finalized only when the file is closed, so code that uses fsync as a durability point must treat close as that point instead. Second, streaming only works for new, sequential writes. Modifying an existing file or writing out of order makes gcsfuse fall back to the older staged path: the whole file is written to a temporary file under file-system:temp-dir (default /tmp) and uploaded on close or fsync. A workload that seeks while writing therefore needs local disk the size of its largest open file.

Running on GKE

On GKE, the Cloud Storage FUSE CSI driver runs gcsfuse as a sidecar container injected into your pod, and mounts buckets as ephemeral or persistent volumes. Enable injection with the pod annotation gke-gcsfuse/volumes: "true", and authenticate through Workload Identity by giving the pod's Kubernetes service account access to the bucket. Mount options use colon-separated keys, such as file-cache:max-size-mb:-1.

apiVersion: v1
kind: Pod
metadata:
  name: trainer
  annotations:
    gke-gcsfuse/volumes: "true"            # inject the gcsfuse sidecar
    gke-gcsfuse/memory-limit: "0"          # no sidecar limit; size the node instead
spec:
  serviceAccountName: trainer-ksa          # bound to a Google service account with bucket access
  containers:
  - name: train
    image: us-docker.pkg.dev/acme/ml/train:1.4
    volumeMounts:
    - {name: data, mountPath: /data, readOnly: true}
    - {name: ckpt, mountPath: /ckpt}
  volumes:
  - name: gke-gcsfuse-cache                # custom cache volume for the file cache
    emptyDir: {}                           # on nodes with local SSD ephemeral storage
  - name: data
    csi:
      driver: gcsfuse.csi.storage.gke.io
      readOnly: true
      volumeAttributes:
        bucketName: acme-ml-data
        mountOptions: "implicit-dirs,file-cache:max-size-mb:-1,metadata-cache:ttl-secs:-1,metadata-cache:stat-cache-max-size-mb:-1,metadata-cache:type-cache-max-size-mb:-1"
  - name: ckpt
    csi:
      driver: gcsfuse.csi.storage.gke.io
      volumeAttributes:
        bucketName: acme-ml-checkpoints

The sidecar has resource limits of its own, set with gke-gcsfuse/cpu-limit, gke-gcsfuse/memory-limit and gke-gcsfuse/ephemeral-storage-limit annotations. Google's performance guide suggests "0", no limit, for demanding workloads so the sidecar is not throttled or killed. The file cache lives in the pod's ephemeral storage by default; to put it on a specific medium, define a volume named gke-gcsfuse-cache, for example a memory-backed emptyDir. For read-only mounts, the volume attribute gcsfuseMetadataPrefetchOnMount fills the metadata cache at mount time. Cloud Run offers bucket volume mounts built on the same technology; see Cloud Run.

Worked example: training data and checkpoints

A team trains on a 1.2 TB image dataset of 400,000 shard files in a bucket, on nodes with local SSD, and checkpoints to a second bucket. Their first attempt mounted the dataset with default settings. The first epoch was slow, which they expected. Every later epoch was just as slow, and metadata requests stayed high throughout.

Two causes. The 60-second metadata TTL meant every epoch re-fetched attributes for 400,000 files, and without a file cache every read went to Cloud Storage again. They made three changes: an infinite metadata TTL and unlimited stat and type caches, because the dataset is immutable during a run; a file cache on local SSD; and only-dir so the mount exposes only the dataset version being trained on.

Cache size then mattered in a way that surprises people. Epochs read the whole dataset in a shuffled but complete pass, so the access pattern is a cycle over 1.2 TB. If the cache is smaller than the dataset, least-recently-used eviction removes each file shortly before it is needed again, and the hit rate collapses towards zero; the docs warn about this thrashing and advise that the whole dataset fit. They sized the cache at 1.4 TB on a node with enough local SSD, and epochs after the first read from local disk.

Checkpoints go the other way. Each is written once, sequentially, by one rank, which is the pattern streaming writes are built for. The code below treats close, not fsync, as the durability point, and publishes a LATEST marker only after the checkpoint file is complete, so a restarting job never loads a half-written file.

import os, torch

def save_checkpoint(state, step, root="/ckpt/run-42"):
    final = f"{root}/step-{step:07d}.pt"
    with open(final, "wb") as f:        # streaming write: sequential, uploaded as it goes
        torch.save(state, f)
        f.flush()
        os.fsync(f.fileno())            # does NOT finalize the object; close does
    # reaching this line means close succeeded; if it raised, LATEST still names the
    # previous checkpoint, so readers never follow the marker to an incomplete file
    with open(f"{root}/LATEST", "w") as f:
        f.write(os.path.basename(final))

Trade-offs and alternatives

gcsfuse wins when a program needs a file path, data is large and read mostly sequentially, and copying it first would waste time or disk. It loses on metadata-heavy work with many small files, on workloads that need locking, append-heavy logs or in-place updates, and wherever several writers share files. For those, consider copying data to local SSD at job start with gcloud storage cp, a managed file service such as Filestore, or calling the Cloud Storage client libraries directly, which give you explicit control of requests, retries and preconditions. Choosing a storage class for the bucket, covered in Cloud Storage classes, matters too: classes with retrieval fees make repeated epoch reads expensive unless the file cache absorbs them.

Failure modes

  • Invisible directories. Data written by other tools does not appear because implicit-dirs is off.
  • Stale reads. Long metadata TTLs on a bucket that another process updates serve old attributes or content until the TTL expires.
  • Cache thrash or full disk. A cache smaller than a cyclic working set gives no hits; max-size-mb: -1 on a shared boot disk can fill it.
  • ESTALE on close. Two writers on one object; one loses. Give each writer its own object name.
  • fsync assumed durable. With streaming writes, a crash after fsync but before close leaves no finalized object.
  • Staging fallback. Out-of-order writes silently switch to staging in /tmp, which can fill a small root disk.
  • Sidecar starved. On GKE, default sidecar limits throttle or kill gcsfuse under heavy reads, and the application sees I/O errors.

What to do next

  1. List your workload's operations: file sizes, file counts, sequential or random reads, writers per object, and whether data changes during a run.
  2. Mount read-only data with implicit-dirs if needed, only-dir to narrow scope, and long or infinite metadata TTLs when the data is immutable.
  3. Enable the file cache on local SSD or RAM, sized to hold the whole working set, and turn on cache-file-for-range-read for header-first formats.
  4. Treat close as the commit point for writes, write each object from one writer, and publish a marker file last.
  5. On GKE, set sidecar limits deliberately, provide a gke-gcsfuse-cache volume, and use Workload Identity.
  6. Watch Class A and B operation counts and gcsfuse logs after the first run, and move metadata-heavy or multi-writer work off the mount.
Key takeaway: Cloud Storage FUSE gives programs a file path to a bucket, but every uncached call is an object-store request, and several POSIX guarantees are missing. Use it for large, mostly sequential reads and single-writer outputs. Turn on implicit-dirs when other tools wrote the data, lengthen metadata TTLs for immutable data, and size the file cache to hold the whole working set on fast local storage, because a cache smaller than a cyclic scan gives almost no hits. Streaming writes are the default since version 3.0.0, and close, not fsync, finalizes the object. On GKE, give the sidecar enough resources and a cache volume, and keep databases, locks and shared writers off the mount.