Most people meet the Hugging Face Hub through one line: from_pretrained("org/model"). That line hides a version-controlled repository, a resolution step that turns a branch name into a commit, a content-addressed cache on your disk, a storage backend that deduplicates large files, an access-control system, and a decision about whether to execute code someone else wrote. In a notebook none of that matters. In a training pipeline, a CI job or a production service, every one of those layers is a source of slow starts, silent behaviour changes or security exposure.

This article explains the Hub from the bottom up and turns it into operating practice. It covers repos and revisions, the cache layout, downloading and pinning, offline and air-gapped use, publishing models with commits, pull requests and tags, the Xet storage layer, tokens and gated access, and the trust model for weights and remote code. APIs are those of the huggingface_hub library 1.x, which made the hf command the only CLI and changed several defaults; where 1.0 removed something, the article says so.

Advertisement

Repositories, revisions and the resolve step

Every model, dataset and Space on the Hub is a git repository identified by namespace/name. Repos have branches, tags and commits like any git repo, and every file is addressed by a revision. When you ask for a file without a revision, you get main, which is a moving pointer: the owner can push new weights, a changed tokenizer configuration or a different chat template at any moment, and your next cold start will silently pick it up.

The client therefore does two things on every uncached request. It asks the Hub to resolve the revision to a commit sha and a file hash, and it fetches the bytes for that hash, now served from Xet storage through a CDN. Everything about reproducibility follows from separating those steps: resolve once, record the sha, and pass the sha everywhere afterwards.

From repo id to tensors in memory: the Hub, the cache and your processYour codefrom_pretrained / hf_hub_downloadHub APIresolve revision -> commit shaXet storage + CDNchunk-deduplicated file bytesrepo, revfile hashHF_HOME/hub/models--org--name/refs/main -> a1b2c3...v1.2 -> d4e5f6...branch or tag to shasnapshots/<sha>/config.json (link)model.safetensors (link)one folder per commitblobs/<file hash> (bytes)<file hash> (bytes)stored once, sharedbytesHF_HUB_OFFLINE=1resolve from refs and snapshots onlyrevision=<sha>skips the moving branch; reproduciblePin a commit, prefetch into the cache or a local dir, then run offline
A request names a repo and a revision; the Hub resolves it to a commit; the cache stores bytes once under blobs/ and exposes each commit as a snapshot folder of links.

The local cache, explained

The cache lives under HF_HOME, which defaults to ~/.cache/huggingface, in a hub/ directory with one folder per repo, such as models--Qwen--Qwen2.5-0.5B-Instruct. Inside are three subfolders. blobs/ holds file contents named by their hash, so identical files across commits are stored once. snapshots/<sha>/ mirrors the repo's file tree at one commit, with each entry a symlink to a blob. refs/ holds small text files mapping a branch or tag name to the commit sha it last resolved to.

This layout explains several behaviours that confuse people. Downloading a new commit of a 15 GB model where only the tokenizer changed costs only the tokenizer, because the weight blobs are reused. Offline mode works by reading refs/main to find the last known sha and serving the snapshot. On Windows without symlink support, the library falls back to copying files, so the cache uses more disk. And a cache shared between users on a machine needs consistent permissions, or one user's download leaves blobs another cannot read.

Advertisement

Downloading the right way

The library offers two primitives. hf_hub_download fetches one file and returns its local path; snapshot_download fetches a whole revision, optionally filtered, and returns the snapshot folder. Model classes in Transformers call these for you, but calling them directly gives you control over the revision, the file set and when the network is touched.

from huggingface_hub import HfApi, hf_hub_download, snapshot_download

REPO = "Qwen/Qwen2.5-0.5B-Instruct"

# 1. Resolve the moving branch to an immutable commit once, and record it in config.
sha = HfApi().model_info(REPO, revision="main").sha

# 2. Fetch only what inference needs, at that commit.
local = snapshot_download(
    REPO,
    revision=sha,
    allow_patterns=["*.safetensors", "*.json", "tokenizer*", "*.txt"],
)

# 3. One file is enough for a config check or a small asset.
cfg_path = hf_hub_download(REPO, "config.json", revision=sha)

# 4. Load from the snapshot folder; nothing else is resolved over the network.
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained(local)
model = AutoModelForCausalLM.from_pretrained(local, dtype="auto")

The filter matters more than it looks. Many repos carry the same weights in several formats, such as safetensors, legacy PyTorch .bin files, GGUF or ONNX exports, and an unfiltered snapshot can download several times what you need. allow_patterns limits the fetch to the formats you load. The same pattern works from the shell with the hf CLI:

# huggingface_hub 1.x: the CLI is "hf"; huggingface-cli was removed in 1.0
hf auth login                                   # interactive; in CI set HF_TOKEN instead
hf download Qwen/Qwen2.5-0.5B-Instruct --revision "$MODEL_SHA" \
    --include "*.safetensors" "*.json" "tokenizer*" --local-dir ./model
hf upload my-org/churn-classifier ./out . --commit-message "epoch 3, val_f1 0.871"
hf cache ls                                     # what is in the cache, how big, when last used
hf cache prune --dry-run                        # unreferenced revisions that would be deleted

If you are migrating from 0.x, note what 1.0 removed: the huggingface-cli command, the resume_download and local_dir_use_symlinks parameters, the use_auth_token parameter in favour of token, and the Repository git wrapper in favour of the HTTP commit API. The old hf cache scan and hf cache delete commands became hf cache ls, rm and prune. The library also moved its HTTP layer to httpx, so proxies are configured with the standard HTTPS_PROXY environment variables.

Offline, air-gapped and reproducible

Production inference should not depend on the Hub being reachable at start-up. Setting HF_HUB_OFFLINE=1 makes the library resolve everything from the local cache and fail fast instead of waiting on network timeouts. Combine it with a build step that fetches a pinned sha into the image or a volume, and the service becomes reproducible and independent of the Hub at runtime.

# Build stage: fetch with a build secret, never an ENV token baked into a layer.
FROM python:3.12-slim AS fetch
RUN pip install --no-cache-dir "huggingface_hub>=1.0,<2"
ARG MODEL_SHA
RUN --mount=type=secret,id=hf_token \
    HF_TOKEN="$(cat /run/secrets/hf_token)" \
    hf download my-org/churn-classifier --revision "$MODEL_SHA" --local-dir /model

FROM my-serving-base:2026.10
COPY --from=fetch /model /model
ENV HF_HUB_OFFLINE=1
CMD ["serve", "--model", "/model"]

Two choices in that file are deliberate. The token arrives as a build secret, so it is never written into an image layer; an ENV HF_TOKEN line would publish it to everyone who can pull the image. And --local-dir produces a plain folder of files rather than the cache layout, which is easier to copy between stages and to inspect. For fleets that cannot reach the internet at all, run the fetch stage on a connected builder and promote the resulting artefact, or mirror approved repos into an internal store and point the fleet there.

Publishing: commits, pull requests and tags

The Hub is also where you publish your own fine-tunes and datasets. With 1.x, uploads go through the HTTP commit API rather than a local git clone, which avoids cloning multi-gigabyte repos just to add a file.

from huggingface_hub import HfApi

api = HfApi()                                   # token from HF_TOKEN or a prior login
repo = "my-org/churn-classifier"
api.create_repo(repo, private=True, exist_ok=True)

api.upload_folder(
    repo_id=repo,
    folder_path="out/",
    commit_message="epoch 3, val_f1 0.871",
    ignore_patterns=["optimizer*", "*.pt", "checkpoint-*/"],   # ship weights, not training state
)
sha = api.model_info(repo).sha                  # the commit you just created
api.create_tag(repo, tag="v1.2.0", revision=sha, tag_message="promoted after eval")

# Changes to a model other people depend on: open a pull request instead of pushing to main.
api.upload_file(path_or_fileobj="README.md", path_in_repo="README.md",
                repo_id=repo, create_pr=True, commit_message="Document eval set")

Treat a model repo like a release branch. Push training output to a private repo, upload only the artefacts consumers need, and exclude optimizer state and intermediate checkpoints, which can be larger than the weights. Promote a specific commit with a tag such as v1.2.0 after evaluation, and have consumers pin either the tag or, better, the commit sha the tag pointed to when they adopted it, because tags can be moved. Changes to repos other teams depend on should go through create_pr=True, which opens a Hub pull request that can be reviewed before it lands on main. The model card, the repo's README.md, is where training data, evaluation results and intended use belong.

Xet storage: why uploads got cheaper

Large files on the Hub used to be stored with Git LFS, where changing any byte of a file meant uploading the whole file again. The Hub has migrated its repositories to Xet, a storage backend that splits files into content-defined chunks and deduplicates at chunk level. When you re-upload a fine-tuned checkpoint that shares most of its structure with the previous one, or append rows to a large dataset file, only the changed chunks are transferred. The hf_xet client is the default transfer path in huggingface_hub 1.x; the older hf_transfer package is no longer supported, the HF_HUB_ENABLE_HF_TRANSFER variable is ignored, and HF_XET_HIGH_PERFORMANCE is the documented switch for high-bandwidth machines.

The practical consequence is that frequent checkpoint uploads during training are far less wasteful than they were, so pushing every evaluated epoch to a private repo is a reasonable experiment-tracking strategy. Deduplication helps the uploader; a fresh download still needs every chunk.

Tokens, private repos and gated models

Access is controlled with user access tokens. Fine-grained tokens, scoped to specific repos or organisations and to read or write, are the recommended kind; 1.0 removed the older notion of a generic write token from the library's login flow. The library reads the token from HF_TOKEN or from the file hf auth login writes, and every download and upload function also accepts token= directly.

Gated models are public repos whose owner requires users to accept terms or be approved before downloading. A gated download needs a token belonging to an account that has been granted access; an anonymous request or a token from another account fails with an authorisation error even though the repo page is visible. In CI, use a dedicated machine account with read-only, repo-scoped tokens, store the token in the CI secret store, and rotate it on a schedule. Never put a write token on an inference host.

The trust model: weights are code until proven otherwise

Two Hub features can execute code on your machine. The first is pickle: legacy PyTorch .bin checkpoints are pickle files, and unpickling can run arbitrary Python. Since PyTorch 2.6, torch.load defaults to weights_only=True, which restricts what can be unpickled, but older stacks and other loaders do not. Safetensors files contain only a JSON header and raw tensor bytes, so loading them cannot execute code; prefer them and filter downloads to them. The Hub scans uploaded pickles and flags suspicious imports, which is useful but not a guarantee.

The second is trust_remote_code=True, which tells Transformers to import Python modules from the repo to define a custom architecture or tokenizer. That code runs with your process's permissions, and it changes whenever the repo's main changes. If you must use it, pin the commit sha, read the code at that sha, and treat an upgrade as a code review. These risks are a model-shaped version of the problems in software supply chain security, and the same controls apply: pinning, allowlists, mirrors and provenance.

Failure modes

SymptomCauseFix
Model output changed with no deployLoading main, which the owner updatedPin a commit sha; record it with the deployment
Slow or rate-limited starts in CIEvery runner downloads the same filesShared cache volume or a prebuilt artefact; HF_HUB_OFFLINE=1 at runtime
Disk full on training nodesCache keeps every revision and formatallow_patterns; hf cache ls and hf cache prune on a schedule
Authorisation error on a public-looking repoGated model, token lacks accessRequest access with the machine account; pass its token
huggingface-cli: command not foundhuggingface_hub 1.x removed itUse hf; pin the library version in images
Token leaked in an imageENV HF_TOKEN in a DockerfileBuild secrets; rotate the token immediately
Unexpected code ran at loadPickle weights or trust_remote_code on a moving revisionSafetensors only; pin and review remote code

What to do next

  1. Find every from_pretrained and download call in your services and record which revision each resolves.
  2. Replace branch names with commit shas in configuration, and log the sha at start-up.
  3. Add allow_patterns or --include filters so only safetensors and configs are fetched.
  4. Move downloads into the image build with a build secret, and set HF_HUB_OFFLINE=1 in production.
  5. Upgrade scripts to huggingface_hub 1.x: replace huggingface-cli, use_auth_token and the Repository class.
  6. Create a machine account with fine-grained, read-only tokens for CI and inference, and rotate them.
  7. Audit for trust_remote_code=True and legacy .bin weights; pin or convert each one.
  8. Continue with the Transformers library architecture, pipelines, datasets and model registries for promotion workflows.
Key takeaway: The Hugging Face Hub is a git-backed store addressed by repo and revision, with a content-addressed local cache under HF_HOME. Resolve main to a commit sha once and pin it everywhere, fetch only safetensors and configs, download at build time with a secret, and run with HF_HUB_OFFLINE=1. Publish through commits, pull requests and tags with huggingface_hub 1.x and the hf CLI, use fine-grained tokens, and treat pickle weights and trust_remote_code as code to review.