A credential committed to a repository does not stay in one place. It is copied into every clone, every fork, CI caches, build logs, container layers and code search indexes, and deleting the line in a later commit removes it from none of them. Secret scanning is the set of detectors and checkpoints that find these credentials, ideally before they leave a laptop and otherwise as soon after as possible, and the workflow that turns a finding into a revoked key.

This article explains how a detector decides that a string is a secret, where to place scans so each one catches what the others miss, how to configure gitleaks and TruffleHog for those places, how to adopt scanning on a repository with years of history, and how to remediate a real leak in the right order.

Advertisement

What scanning is for, and what it is not

Scanning is detection. It does not make a leaked key safe, and it does not replace the real fix, which is to have fewer long-lived secrets at all: workload identity, short-lived tokens issued at runtime and a secrets manager that injects values rather than files in the repository. That architecture is covered in secrets management architecture. Scanning is the net under it, because people will still paste a token into a test fixture at the end of a long day.

Think of a lifecycle. A secret is created at an issuer (AWS, GitHub, Stripe, your own auth service), copied into a developer environment, and may then cross boundaries: staged, committed, pushed, merged, forked. Every boundary it crosses multiplies the places it lives and the effort to contain it. The goal of the architecture is to catch the secret at the earliest boundary, and to make sure that whatever is caught late is still revoked quickly.

Four checkpoints, each with a different job

Four checkpoints for a secret, from keyboard to history1. Pre-commit hookstaged diff, seconds2. Push protectionserver-side, on push3. CI on the PRcommit range, minutes4. Scheduled scanall history, all reposDetector pipelinekeyword prefilter, regex with secret group, entropy floor, allowlists, optional live verificationFindings storefingerprint and secret hash, never the secretTriage and routingowner, severity, SLA, dedupRevoke and rotateat the issuer, firstInvestigate useissuer access logsClean up and closehistory rewrite is optional
Each checkpoint feeds the same detector pipeline. Findings are stored by fingerprint and hash and routed to an owner; remediation starts with revocation at the issuer.

No single checkpoint is enough, because each can be skipped or sees only part of the picture. A pre-commit hook runs on the developer machine and sees the staged diff before a commit exists, which is the cheapest moment to fix anything; but hooks are opt-in and trivially skipped with git commit --no-verify. Server-side push protection, such as GitHub's, runs when commits reach the hosting service and cannot be skipped by the client, but it only blocks pattern types the provider supports and it can be bypassed with a stated reason. A CI job on the pull request scans the exact commit range under review with your own rules. A scheduled scan walks all history in all repositories, including branches nobody has looked at in years, and catches what predates the other three.

CheckpointSeesCost of a findingCan be skipped
Pre-commit hookStaged changes onlyEdit the file, commit againYes, by the developer
Push protectionCommits being pushedAmend before the push landsWith a recorded bypass reason
CI on the PRCommit range of the PRSecret is on the server: rotateOnly by changing CI config
Scheduled scanFull history, all refs, all reposSecret may be old and copied: rotate and investigateNo

CI logs, container images, tickets and chat exports leak credentials too; the same detectors can read them through a file or stdin mode.

Advertisement

How a detector decides that a string is a secret

A detector is a small pipeline that runs per line or per diff hunk. First a keyword prefilter: the rule declares cheap literal strings, such as AKIA or xoxb-, and the expensive regex runs only on content that contains one. Then a regular expression with a capture group that isolates the secret itself from its surroundings, so the finding reports the token rather than the whole assignment. Then an entropy floor that discards low-randomness matches such as password = changeme. Then allowlists for paths, commits and stopwords. Finally, for some tools, live verification against the issuer.

The strongest signal by far is a structured prefix. Many issuers now put a fixed, documented prefix on their tokens precisely so scanners can find them: AWS access key IDs begin with AKIA for long-term keys and ASIA for temporary ones, GitHub personal access tokens begin with ghp_ or github_pat_, Slack bot tokens with xoxb- and Stripe live secret keys with sk_live_. A prefixed token rule has very few false positives. Generic rules, which look for anything shaped like a key next to a word like secret or token, are where the noise comes from.

Entropy, worked through, and why it misfires

Shannon entropy measures how unpredictable the characters of a string are: H = -sum over distinct characters of p(c) log2 p(c), where p(c) is the character's frequency in the string. It is measured in bits per character. A string of one repeated character scores 0. A hex string can score at most 4 bits per character because it has 16 symbols; base64 can reach 6. Scanners use it as a floor: a match whose entropy is below, say, 3.5 is treated as a placeholder rather than a secret.

import math
from collections import Counter

def shannon(s: str) -> float:
    n = len(s)
    return -sum(k / n * math.log2(k / n) for k in Counter(s).values())

The values below were computed with that function. The two AWS strings are the example credentials from AWS documentation, not real keys.

String (length)Bits per charIs it a secret?
wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEY (40)4.66Yes: secret access key shape
DATABASE_URL_PRIMARY_REPLICA_HOSTNAME_01 (40)4.03No: an identifier
0123456789abcdef0123456789abcdef01234567 (40)3.97No: a sequence
AKIAIOSFODNN7EXAMPLE (20)3.68Yes: access key ID shape
3f786850e387550fdab836ed7e6dc881de23001b (40)3.63No: a SHA-1 of a file
correcthorsebatterystaple (25)3.36Possibly: a real passphrase

Read the table carefully. A long constant name scores higher than a real access key ID, a git commit hash looks exactly like a hex token, and a human passphrase scores lower than all of them. Entropy separates random from repetitive; it does not separate secret from public. That is why good rules combine a prefix or context regex with entropy, and why a pure entropy scan of a repository full of lockfiles and hashes produces thousands of findings nobody will triage.

Verification: the strongest signal, with a cost

TruffleHog adds a step after detection: for many credential types it calls the issuer's API with the candidate and classifies the result as verified (the credential works), unverified (detected but not confirmed) or unknown (the verification attempt itself failed). Filtering to verified results gives a very low false-positive rate, because a live credential is the thing you care about most.

Two cautions. Verification sends the candidate to a third party from wherever the scanner runs, so decide whether that is acceptable and allow the egress explicitly. And unverified does not mean safe: a key that fails verification because of a network error is still a key. Page on verified, ticket on unverified and unknown.

Configuration at each checkpoint

Gitleaks reads rules from a TOML file. Since version 8.19 its subcommands are git (scans commits through git log -p), dir and stdin; the older detect and protect commands are hidden and deprecated. A custom rule for an internal token format looks like this:

# .gitleaks.toml
[extend]
useDefault = true

[[rules]]
id = "acme-service-token"
description = "Acme internal service token"
regex = '''acme_svc_([A-Za-z0-9]{32})'''
secretGroup = 1
entropy = 3.5
keywords = ["acme_svc_"]

[[allowlists]]
description = "Vendored fixtures and docs examples"
paths = ['''^testdata/''', '''^docs/examples/''']

The pre-commit hook published by the project runs gitleaks git --pre-commit --redact --staged --verbose, and CI should scan only the pull request's range and emit SARIF so findings appear in the code review UI:

# .pre-commit-config.yaml
repos:
  - repo: https://github.com/gitleaks/gitleaks
    rev: v8.x.y            # pin to a released tag
    hooks:
      - id: gitleaks

# CI step on a pull request
gitleaks git --redact --log-opts="origin/main..HEAD" \
  --report-format sarif --report-path gitleaks.sarif

# Verification-first pass with TruffleHog, failing the job on results
trufflehog git file://. --since-commit main --branch "$PR_BRANCH" \
  --results=verified,unknown --fail

Gitleaks exits 1 when it finds leaks; TruffleHog with --fail exits 183, so wire your CI on those codes rather than parsing output. Always pass --redact or equivalent: a scanner that prints the secret into a CI log has just leaked it a second time.

Adopting scanning on a repository with history

The first full scan of an old repository usually produces hundreds of findings, most of them dead test keys, examples and false positives. If you turn on a blocking CI gate at that moment, every pull request fails for reasons its author did not cause, and the gate is switched off within a week. Instead, separate the backlog from the flow. Run one full-history scan, write its report to a file, and pass it as --baseline-path so subsequent scans report only new findings. Then triage the baseline as a project with an owner and a deadline.

Gitleaks identifies a finding by a fingerprint of the form commit:file:ruleID:line; listing one in .gitleaksignore suppresses exactly that occurrence, and a gitleaks:allow comment suppresses a single line in source. Prefer fingerprints and path allowlists in reviewed files over inline comments. Incremental scans need care too: a range like origin/main..HEAD misses commits on branches that were force-pushed or never merged, which is exactly why the scheduled full scan still exists.

Triage without building a new secret store

A findings pipeline is tempting to build naively: put every match, including the matched value, into a database and a dashboard. That database is now the most valuable credential store in the company. Store a fingerprint, the rule, the location, the commit author and a salted hash of the secret for deduplication, and never the value. The hash lets you notice that the same key appears in five repositories, which changes the severity.

Route each finding to an owner using the same ownership data as code review, typically CODEOWNERS for the path, falling back to the commit author. Severity comes from three questions: is it verified, what can it do (a production cloud key is not a sandbox webhook secret), and how exposed was it (a public repository is assumed compromised within minutes, an internal one is not). Give each severity a response time and measure time from detection to revocation, because that, not the count of findings, is the number that reflects risk.

Remediation in the right order: a worked example

At 10:02 a developer pushes a feature branch to a public repository containing an AWS access key in a config file. Push protection is not enabled on that repository, and the CI job on the pull request reports a verified finding at 10:09. Deleting the line and force-pushing is the wrong response: anyone watching public pushes may already have the key. The right order is:

  1. Revoke or deactivate the key at the issuer, and issue a replacement through the normal secrets path. Revocation is the only step that ends the exposure.
  2. Deploy the replacement and confirm that the old key is rejected, so nothing in production still depends on it.
  3. Investigate use between 10:02 and revocation in the issuer's audit logs, for AWS CloudTrail entries made with that access key ID, and escalate to incident response if anything unexpected appears.
  4. Remove the value from the current tree and move the code to read it from the secrets manager. Rewriting history is optional and mostly cosmetic once the key is dead; do it only if the repository policy requires it.
  5. Close the finding with the time to revoke recorded, and add a rule or a pre-commit hook if this pattern was not caught earlier than CI.

Rotation is much easier when the application already supports two valid credentials at once; secret rotation in Kubernetes walks through that overlap window.

Failure modes

  • The scanner leaks the secret. Logs, SARIF files and chat alerts that include the matched value spread it further. Redact everywhere.
  • Alert fatigue. Generic high-entropy rules across lockfiles and fixtures bury real findings. Start with prefixed rules and verified results, then add generic rules per path.
  • Bypass as a habit. Push protection asks for a reason when someone pushes through a block. Review bypasses weekly; a cluster from one team usually means a missing test fixture convention, not malice.
  • Blind spots. Secrets that are base64-encoded, split across lines, inside notebook outputs or in binary files evade line-based regexes. Scan notebooks and images deliberately.
  • Only the default branch is scanned. Abandoned branches and tags keep old commits reachable. Scheduled scans must walk all refs.

Trade-offs to decide explicitly

Block or warn: blocking at push and in CI stops leaks but costs developer time on false positives, so block only on high-precision rules and warn on the rest. Verification or not: it sharpens precision but requires outbound calls with candidate credentials. Hosted or self-run: provider push protection is hardest to skip, while self-run tools cover custom formats; most organisations need both. Treat scanning as one control inside a broader software supply chain and application security testing programme.

What to do next

  1. Turn on server-side push protection for every repository where your hosting provider supports it.
  2. Publish a pre-commit configuration with gitleaks and make it part of the repository template.
  3. Add a CI job that scans the pull request range with redaction and SARIF output, failing on prefixed and verified findings.
  4. Run one full-history scan per repository, save it as a baseline, and give the backlog an owner and a deadline.
  5. Build the findings store with fingerprints and salted hashes only, routed by code ownership.
  6. Write the revocation runbook for your top five credential types and measure detection-to-revocation time.
  7. Replace the credentials you find most often with short-lived, identity-based access so there is less to leak.
Key takeaway: Secret scanning finds credentials at four checkpoints: the pre-commit hook, server-side push protection, CI on the pull request and scheduled scans of all history. Detectors combine prefixes, regex, entropy, allowlists and sometimes live verification, and entropy on its own cannot tell a secret from a hash. Adopt it on old repositories with a baseline, store fingerprints rather than secrets, route findings by ownership, and remediate by revoking at the issuer first. Time from detection to revocation is the metric that matters.