Filestore is Google Cloud's managed NFS service. You ask for a capacity, a tier and a network; Google runs the file server in a network it manages and gives you an IP address and an export path. Any Linux client in your VPC mounts it with the ordinary kernel NFS client, and every client sees the same POSIX file system: directories, renames, permissions and locks, with no object-store semantics to work around.

That makes Filestore the answer to a narrow but common need: several machines must read and write the same files, and the application expects a real file system. Typical cases are shared home and project directories, content management systems, legacy applications lifted from on-prem NAS, GKE workloads that need ReadWriteMany volumes, and training jobs that share datasets and checkpoints. This page explains how it works from the network up, how to pick a tier and size it, how to mount it for full performance, how to use it from GKE, how to protect the data, and what goes wrong. Tier limits and flags were checked against Google's Filestore documentation on 2026-10-04.

How Filestore works

NFS is a client-server protocol. The server owns the file system and its metadata. Clients send operations over the network: look up a name, read 512 KiB at an offset, write, rename, get attributes. The Linux client caches attributes and data and decides when to revalidate them, which gives "close-to-open" consistency. When one client closes a file, another client that opens it afterwards sees the writes. Two clients writing the same byte range at once without locks get no guarantees, the same as on any NFS server.

Filestore runs that server for you. You never get a VM or a disk to manage. You choose capacity, and on most tiers a performance level, and Google handles the storage behind it, patches and failure handling. Compare this with Cloud Storage FUSE, which presents a bucket as files. FUSE is cheap and scales to any size, but renames are copies and there is no file locking across clients. Filestore costs more per GiB and has a capacity ceiling per instance, but it behaves like the file server applications were written for.

Filestore: an NFS server in a Google-managed network, reached over a private rangeYour VPCCompute Engine VMmount -o nconnectGKE nodeCSI driver + PVCBatch / training VMsmany readersOn-prem via VPNneeds route exportReserved range/29 basic, /26 zonal and regionalGoogle-managed producer networkFilestore instanceIP 10.x.x.2, share /dataSnapshotsinside instanceCapacity1 GiB or 256 GiB stepsRegional tier: replicas across zonesNFSv3 / v4.1peering or PSABackupsseparate copy in a regionbackupClients see an IP and an export path.No agent, no FUSE, standard kernel NFS client.Peering is not transitive: networks peered to your VPC, and on-prem networks, need routes to the reserved range.
A Filestore instance lives in a Google-managed producer network connected to your VPC over a reserved range. Clients mount it over NFS. Snapshots live inside the instance; backups are separate copies. The regional tier keeps replicas in several zones.

Choosing a tier

Tiers differ in availability, protocol, how capacity can change and how performance is set. The current tiers in Google's documentation:

TierCapacityResizeAvailabilityNFSPerformance
Basic HDD1 to 63.9 TiBUp only, 1 GiB stepsZonalv3Fixed; about 100 MiB/s below 10 TiB
Basic SSD2.5 to 63.9 TiBUp only, 1 GiB stepsZonalv3Fixed; up to 1,200 MiB/s read, 350 MiB/s write
Zonal1 to 9.75 TiB, or 10 to 100 TiBUp or down, 256 GiB or 2.5 TiB stepsZonalv3, v4.1Configurable IOPS per TiB
Regional100 GiB to 9.75 TiB, or 10 to 100 TiBUp or downRegionalv3, v4.1Configurable IOPS per TiB
Enterprise multishares for GKE1 to 10 TiB per instanceUp or down, 256 GiB stepsRegionalv3, v4.1Scales with capacity

Three rules fall out of that table. First, basic tiers can only grow, so over-provisioning a basic instance is permanent until you migrate. Second, basic tiers support backups but not snapshots or replication; Zonal and Regional support all three. Third, only Regional survives the loss of a zone. A Zonal instance in a zone that goes down is unavailable until the zone recovers, however many clients you have elsewhere.

On Zonal and Regional, performance is set as provisioned IOPS per TiB, and custom performance is on by default. The documented ranges are 4,000 to 17,000 IOPS per TiB for the lower capacity band and 3,000 to 7,500 for the 10 to 100 TiB band. Throughput rises with IOPS, but take the exact figure for your configuration from the performance page instead of deriving it.

Networking: reserved ranges and connect modes

Filestore is not inside your VPC. The instance lives in a Google-managed producer network, connected to yours by VPC Network Peering or, with PRIVATE_SERVICE_CONNECT, a Private Service Connect endpoint. For peering there are two connect modes. DIRECT_PEERING peers directly and takes an IP range for the instance. PRIVATE_SERVICE_ACCESS uses an allocated range you have already reserved for Google services, and it is required when the instance attaches to a Shared VPC. Basic instances need a /29 block; Zonal, Regional and Enterprise need a /26. Plan these ranges like any subnet. An overlap with an on-prem range is painful to undo.

Peering is not transitive. Clients in your VPC reach the instance, but clients in a third network peered with your VPC do not. On-prem clients over Cloud VPN or Interconnect reach it only if your Cloud Router advertises the reserved range to on-prem. Most "mount hangs from the data centre" tickets come down to that missing route. If you restrict egress with firewall rules or policies, allow the NFS ports listed in the Filestore firewall documentation from client subnets to the reserved range.

Creating and mounting an instance

Creating a Zonal instance with a share called data, export rules and deletion protection. First, exports.json says who may mount and how root is treated:

{
  "--file-share": {
    "capacity": "8192",
    "name": "data",
    "nfs-export-options": [
      { "access-mode": "READ_WRITE", "ip-ranges": ["10.20.0.0/20"], "squash-mode": "ROOT_SQUASH",
        "anon_uid": 1003, "anon_gid": 1003 },
      { "access-mode": "READ_ONLY",  "ip-ranges": ["10.30.0.0/20"], "squash-mode": "ROOT_SQUASH" }
    ]
  }
}
gcloud filestore instances create ml-shared \
    --zone=us-central1-a \
    --tier=ZONAL \
    --performance=max-iops-per-tb=8000 \
    --network=name=prod-vpc,connect-mode=PRIVATE_SERVICE_ACCESS \
    --flags-file=exports.json \
    --deletion-protection

access-mode is READ_WRITE or READ_ONLY. squash-mode is ROOT_SQUASH or NO_ROOT_SQUASH; anon_uid and anon_gid are only valid with ROOT_SQUASH and default to 65534. Squash root unless a client must manage ownership, because a root user on any allowed client otherwise has root on the share. Then mount it on a client:

sudo apt-get install -y nfs-common
sudo mkdir -p /mnt/data
sudo mount -t nfs -o rw,hard,async,nconnect=2,rsize=524288,wsize=524288 \
    10.20.64.2:/data /mnt/data
# /etc/fstab, so it survives reboots and waits for the network:
# 10.20.64.2:/data /mnt/data nfs rw,hard,nconnect=2,rsize=524288,wsize=524288,_netdev 0 0

Getting the performance you paid for

Most Filestore performance complaints are client-side. Google's guidance is concrete. Use nconnect=2 for instances of 1 to 9.75 TiB and nconnect=7 for 10 to 100 TiB and the basic tiers; nconnect opens several TCP connections per mount, so one client is not limited to one connection's throughput. Use rsize and wsize of 524288, or 1048576 on basic tiers. Use hard mounts with async, because sync turns every write into a round trip. Set read_ahead_kb to about 20 MB for large sequential reads.

Two more limits come from the clients themselves. A single VM's throughput is capped by its own network egress and CPU, so Google notes that at least four client VMs are needed to reach an instance's full performance. And small-file workloads are bound by metadata round trips, not bandwidth. Unpacking a million-file archive onto NFS is slow on any tier. Ship one archive, unpack it on local SSD, or keep datasets in large shard files.

# Measure before you tune: sequential read and 4 KiB random read
fio --name=seq --directory=/mnt/data/bench --rw=read --bs=1M --size=8G \
    --numjobs=8 --ioengine=libaio --direct=1 --group_reporting
fio --name=rand --directory=/mnt/data/bench --rw=randread --bs=4k --size=2G \
    --numjobs=16 --iodepth=32 --ioengine=libaio --direct=1 --group_reporting

Filestore volumes on GKE

On GKE, the Filestore CSI driver provisions instances from PersistentVolumeClaims. It is on by default in Autopilot; on Standard clusters enable it with gcloud container clusters update CLUSTER --update-addons=GcpFilestoreCsiDriver=ENABLED. GKE installs StorageClasses for each tier: standard-rwx (Basic HDD), premium-rwx (Basic SSD), zonal-rwx, regional-rwx, enterprise-rwx and enterprise-multishare-rwx. For a Shared VPC cluster, define your own class with the host network and connect-mode:

apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
  name: filestore-zonal-shared
provisioner: filestore.csi.storage.gke.io
volumeBindingMode: WaitForFirstConsumer
allowVolumeExpansion: true
parameters:
  tier: zonal
  network: projects/host-project/global/networks/prod-vpc
  connect-mode: PRIVATE_SERVICE_ACCESS
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: datasets
spec:
  accessModes: ["ReadWriteMany"]
  storageClassName: filestore-zonal-shared
  resources:
    requests:
      storage: 2Ti

Each PVC on a normal class creates a whole instance, with that tier's minimum capacity and a reserved range. A namespace with twenty small PVCs on standard-rwx creates twenty 1 TiB instances. The multishare class packs many smaller shares into one Enterprise instance, which is the fix when many teams each need a modest RWX volume. GKE for ML, in depth covers how shared volumes fit next to accelerator node pools.

Snapshots, backups and replication

Three mechanisms protect data, and they answer different questions. Snapshots (Zonal, Regional, Enterprise) are point-in-time copies inside the instance. You can revert to one quickly, but they share the instance's fate. Backups (all tiers) are separate copies stored in a region, which you can restore to a new instance even if the original is deleted. Replication (Zonal, Regional, Enterprise) keeps a copy on another instance for disaster recovery.

A sound default is daily snapshots kept for a week for "I deleted a directory" recovery, plus daily backups kept for 30 days or more for "the instance is gone" recovery. --deletion-protection stops an accidental instance delete. Test a restore before you need one: restoring a backup to a new instance gives it a new IP, so clients must remount, and the fstab entries or the PersistentVolume must change too.

Worked example: a shared dataset and checkpoint store

An ML team has a 5 TiB image dataset and writes checkpoints of about 200 GiB per run. Sixteen training VMs read the dataset at the start of each epoch, and 2 to 4 runs execute at once. They want one shared path and can tolerate a zonal outage, since the jobs can restart.

Capacity: 5 TiB of data, plus four runs times three retained checkpoints times 200 GiB (about 2.3 TiB), plus 15% free space, comes to roughly 8.5 TiB. That fits the Zonal lower band (1 to 9.75 TiB), which can shrink later. Performance: at 8,000 IOPS per TiB, 8.5 TiB gives about 68,000 IOPS. With the dataset packed into large shard files, the real constraint is sequential throughput across sixteen readers, so run the fio sequential job from four VMs at once and raise max-iops-per-tb if the aggregate falls short. Mount with nconnect=2 and rsize=524288, put the dataset in large shard files, and keep scratch data on local SSD. Snapshots daily, backups weekly, and the dataset's canonical copy in Cloud Storage, so Filestore holds a working copy rather than the only one.

Failure modes

  • Full file system. Writes fail with ENOSPC and applications fail in odd ways. Alert at 80% used; resizing is online, but on basic tiers only upward.
  • Hung mounts. A hard mount blocks I/O while the server is unreachable instead of returning errors. That is correct for data safety, but processes appear frozen during a zonal outage or a network change. Do not switch to soft mounts to hide it; soft can turn a timeout into silent data loss for writers.
  • Permission surprises. With root squash, files created as root on a client belong to the anonymous UID. Containers running as root then cannot read files created by other pods. Run as a fixed non-root UID, or set fsGroup.
  • Unreachable from peered or on-prem networks. Peering is not transitive; add route advertisements for the reserved range.
  • PVC sprawl on GKE. Each claim on a non-multishare class is an instance with a minimum size. Audit Filestore instances created by the CSI driver monthly.
  • Slow small files. Metadata latency dominates; batch files into shards or use local SSD for scratch.

Trade-offs

OptionPick it whenAvoid it when
Filestore BasicLow-cost shared files, a zonal outage is acceptableYou may need to shrink, or need snapshots
Filestore ZonalShared files with tunable performance and snapshotsA zone outage must not stop the workload
Filestore RegionalShared state that must survive a zone lossData is reproducible and cost matters most
Cloud Storage FUSELarge read-mostly datasets, petabyte scaleApps need locks, atomic renames or in-place updates
Persistent DiskOne writer per disk; databasesSeveral nodes must write the same files

What to do next

  1. Write down who needs shared access, how much capacity, and whether a zonal outage is acceptable. That picks Basic, Zonal or Regional.
  2. Reserve a /26 (or /29 for basic) that cannot collide with on-prem ranges, and decide on PRIVATE_SERVICE_ACCESS if you use Shared VPC.
  3. Create the instance with export rules, root squash and --deletion-protection.
  4. Mount from four test clients with the recommended options and run the fio jobs; record the results next to the tier settings.
  5. On GKE, use one StorageClass per tier with the right network, and prefer multishare when many teams need small RWX volumes.
  6. Schedule snapshots and backups, then restore a backup to a new instance and remount a client to prove the procedure.
  7. Add alerts on capacity used, and check route advertisement before any on-prem client tries to mount. For private access to Google APIs from the same clients, see Private Google Access and Private Service Connect.
Key takeaway: Filestore is a managed NFS server reached over a peered private range: pick Basic for cheap zonal shares that only grow, Zonal for tunable performance and snapshots, and Regional when the share must survive a zone loss. Reserve ranges carefully, remember peering is not transitive, mount with nconnect and large rsize and wsize from several clients, squash root, watch PVC sprawl on GKE, and prove a backup restore before you depend on it.