Perspective API is Jigsaw's hosted toxicity classifier. You send it a piece of text and it returns, for each attribute you ask about, a number between 0 and 1. For close to a decade it was the default way for news sites and forums to triage comments, and for language-model researchers it became the default ruler for measuring how toxic a model's output is. Both uses are now on a clock: the service is being shut down, and Jigsaw states it will remain active until 31 December 2026, with no extension and no direct migration support.

That makes this a useful moment to understand the system properly, because the work in front of most teams is leaving it without silently changing what their moderation does or what their published numbers mean. This article covers what the score measures, the API contract, a quota-aware client, why research scores are not comparable across time, documented biases, and a replay-and-recalibrate migration, with a worked example and a checklist.

Status as of October 2026

The facts as published on the Perspective site, as of October 2026: the API is being sunset; the service stays active until 31 December 2026; requests for usage and for increased quota were only handled until February 2026; Jigsaw will not offer direct migration support and cannot extend access past the final date. The stated reason is that AI capabilities have moved on and there is less demand for a standalone tool in this area.

Three practical consequences follow. First, if you are below the quota you need, you will not get more, so any migration replay must fit inside your existing quota, which argues for starting now and sampling rather than replaying everything. Second, after the shutdown the API cannot be used to rescore anything, so any score you will ever need for an audit, an appeal or a paper has to be computed and stored before the end of the year. Third, the thresholds embedded in your policy were tuned to this specific model's score distribution and will not transfer to another model unchanged.

What a score means

Jigsaw defines toxicity as a rude, disrespectful or unreasonable comment that is likely to make someone leave a discussion. The models were trained on comments labelled by several human raters each, and the target for a comment was the fraction of raters who called it toxic. So the default score is best read as an estimate of the probability that a reader would perceive the comment as toxic. It is not a measure of how toxic the comment is.

That distinction drives correct use. A mild but unambiguous insult can score higher than a nastier comment phrased politely, because more raters agree on the first. A score of 0.9 means most readers would object, not that the comment is extreme; for severity there is a separate attribute, SEVERE_TOXICITY, trained to be less sensitive to mild profanity. The production attributes are TOXICITY, SEVERE_TOXICITY, IDENTITY_ATTACK, INSULT, PROFANITY and THREAT; other attributes were offered as experimental, and language support varies per attribute.

Two further properties matter. Scores are not calibrated for your community: a gaming forum and a children's homework site have very different base rates and tolerances, so the same 0.7 means different things to each. And the score is a function of the model version, which Jigsaw changed several times, including a 2022 move to multilingual character-level transformer models described in a KDD paper by Lees and colleagues. A threshold is a property of one model version on one traffic mix.

The request and response contract

The service is a single REST method, POST v1alpha1/comments:analyze on commentanalyzer.googleapis.com, authenticated with an API key. The request names the comment, optional context, the languages, the attributes wanted and a few flags:

{
  "comment": {"text": "you are an idiot and nobody wants you here"},
  "languages": ["en"],
  "requestedAttributes": {"TOXICITY": {}, "INSULT": {}, "THREAT": {}},
  "spanAnnotations": true,
  "doNotStore": true
}

The response holds an attributeScores map. Each attribute has a summaryScore for the whole comment and, when span annotations were requested and the model supports them, spanScores with begin and end offsets. Each score carries a value and a type; the default type is PROBABILITY. The response also echoes the languages used to choose a model and the detected languages.

Details that bite: doNotStore asks the service not to retain the text, and you should set it for anything containing user data. Per-attribute scoreThreshold suppresses scores below a value, which saves nothing and loses information, so leave it out and threshold on your side. Span offsets are expressed in UTF-16 code units, not Python string indices, so text with emoji or other astral characters will highlight the wrong substring unless you convert. Requesting an attribute that the chosen language does not support yields an error, not a low score. There is also a comments:suggestscore method for sending corrected labels back as training data; after the shutdown it has no purpose.

A client that behaves under quota

Quota is per project and low by default, and with increase requests closed it is now fixed, so the client must cache, back off and fail in a defined direction. This version scores six attributes, retries only what is retryable, and converts span offsets:

import hashlib, random, time
import requests

URL = "https://commentanalyzer.googleapis.com/v1alpha1/comments:analyze"
ATTRS = ["TOXICITY", "SEVERE_TOXICITY", "IDENTITY_ATTACK", "INSULT", "PROFANITY", "THREAT"]
_cache = {}

def analyze(text, key, lang="en", retries=5):
    k = hashlib.sha256(f"{lang}\x00{text}".encode()).hexdigest()
    if k in _cache:
        return _cache[k]
    body = {"comment": {"text": text}, "languages": [lang],
            "requestedAttributes": {a: {} for a in ATTRS},
            "spanAnnotations": True, "doNotStore": True}
    for attempt in range(retries):
        r = requests.post(URL, params={"key": key}, json=body, timeout=10)
        if r.status_code == 429 or r.status_code >= 500:
            time.sleep(min(30, 2 ** attempt) * random.uniform(0.5, 1.0))
            continue
        r.raise_for_status()          # 4xx such as an unsupported language: do not retry
        data = r.json()["attributeScores"]
        out = {a: v["summaryScore"]["value"] for a, v in data.items()}
        _cache[k] = out
        return out
    raise TimeoutError("perspective unavailable")   # caller decides: hold, never auto-publish

def utf16_span(text, begin, end):
    # Perspective offsets count UTF-16 code units; Python indexes code points.
    units = text.encode("utf-16-le")
    to_py = lambda u: len(units[:2 * u].decode("utf-16-le", errors="ignore"))
    return to_py(begin), to_py(end)

The important line is the last raise. When the scorer is down, a comment platform must choose between publishing unscored text and holding it; write that choice into policy rather than letting an exception handler decide. The cache key includes the language because the same string can route to different models.

Where it sits in a moderation system

A Perspective-backed moderation path, with a shadow scorer for the migrationnew commenttext + languagescore cachehash of textPerspectivecomments:analyzeshadow scorerreplacement, logged onlypolicythresholds per attributemissscoressame textcomparison logboth scores + actionpublishholdhidemoderator decisionsthe labels you will calibrate againstThe policy block owns the thresholds, never the scorer. That is what makes the scorer swappable.
Perspective scores feed a policy layer that owns thresholds and actions; a shadow scorer sees the same traffic during migration.

The architecture that survives a scorer change separates three things. The scorer turns text into numbers. The policy maps numbers to actions such as publish, hold for a moderator or hide, per attribute and per community. The decision log records the text hash, every score, the scorer version, the action and, later, what a human moderator decided. If your thresholds live inside a scorer wrapper, or worse are hard-coded in several services, find them now; they are the part that has to change.

For language-model output rather than user comments, the same split applies, and a page on output filtering for model responses covers the additional problem that a model's reply needs the conversation as context, which Perspective never modelled well. The layered moderation architecture article shows where a scorer like this sits next to policy engines and human escalation, and streaming moderation covers scoring tokens as they are generated.

Perspective as a research metric

In language-model research Perspective became a metric. RealToxicityPrompts (Gehman and colleagues, 2020) released around 100,000 prompts and defined two numbers still widely reported: expected maximum toxicity, the mean over prompts of the highest TOXICITY score among 25 sampled continuations, and toxicity probability, the share of prompts for which at least one of 25 continuations scores 0.5 or more. Detoxification and alignment papers for years reported progress in those units.

The problem is that the ruler changed while people were measuring. Pozzobon and colleagues (2023) rescored published generations with the then-current API and found materially lower toxicity than originally reported, enough to reorder some comparisons between methods. A number scored in 2021 and a number scored in 2024 are measurements from different instruments. With the API gone, even the latest instrument cannot be reproduced.

If you maintain a benchmark or a leaderboard, do three things before the end of 2026: store raw per-generation scores with the scoring date, not just aggregates; rescore your baselines and your current models in one batch so at least one internally consistent set exists; and pick an open, versioned replacement classifier whose weights you can pin, then report both numbers for a transition period. Any table that compares a Perspective number with a number from another classifier is invalid, however similar the definitions sound.

Documented biases

The best documented failure of comment toxicity models, Perspective included, is identity-term bias. Dixon and colleagues (2018), working on Jigsaw's own models, showed that neutral sentences mentioning identities such as gay or Muslim scored higher than the same sentences without them, because those words co-occurred with abuse in the training data. Jigsaw ran the 2019 Unintended Bias in Toxicity Classification competition on this problem. Sap and colleagues (2019) found that tweets in African American English were more likely to be scored as toxic, a dialect effect that a threshold cannot fix.

Later model versions reduced but did not eliminate these effects, and your replacement will have its own. Treat bias as a regression test that runs on every scorer: a template set of neutral and abusive sentences crossed with identity terms and dialect variants, with the per-group false positive rate compared across scorers. The toxicity scanner audit and fairness articles give the measurement method.

Worked example: replay and recalibrate

A community site publishes about 40,000 comments a day. Its policy hides a comment when SEVERE_TOXICITY is at least 0.9, holds it for a moderator when TOXICITY is at least 0.8, and publishes otherwise. The numbers below are illustrative of the procedure, not measurements of any product.

Step one, sample. Inside existing quota, draw a stratified sample of 20,000 recent comments, oversampling the region above 0.5 so the decision boundary is well populated, and include every comment from the last quarter that a moderator acted on. Score each with both Perspective and the candidate replacement, and store both. Step two, match the operating point. At 0.8 the old scorer held 3.1 percent of traffic, so the new hold threshold starts at the score that flags the same share of the same comments. Step three, check agreement, not just rate:

import numpy as np

def matched_threshold(old, new, old_t):
    rate = (old >= old_t).mean()                  # share the current policy flags
    return float(np.quantile(new, 1.0 - rate)), rate

def compare(old, new, old_t, new_t, human):      # human: 1 = moderator removed it
    a, b = old >= old_t, new >= new_t
    overlap = (a & b).sum() / max(1, (a | b).sum())
    prec = lambda f: human[f].mean() if f.any() else float("nan")
    rec = lambda f: f[human == 1].mean()
    return {"jaccard": overlap, "old_precision": prec(a), "new_precision": prec(b),
            "old_recall": rec(a), "new_recall": rec(b)}

Suppose the matched threshold for the replacement is 0.62 and the flag sets overlap with a Jaccard index of 0.68. Equal rates hide real differences: inspect the disagreements in both directions, because those comments tell you what changes for users. If the replacement has higher precision against moderator decisions at equal recall, adopt it; if it misses a category, such as threats phrased without profanity, add a rule or a second classifier for it rather than lowering the global threshold. Step four, run the replacement in shadow on live traffic for two weeks, logging both decisions. Step five, switch the policy to the new thresholds, keep Perspective in shadow until it stops answering, and archive the comparison log as the record of why your moderation changed.

Failure modes

  • Cutover by renaming the scorer. Pointing the old thresholds at a new model changes hold rates overnight, often by multiples, and moderators get flooded or abuse gets through.
  • Fail-open on outage. A timeout path that publishes unscored text becomes the default path on 1 January 2027 if the client is never removed.
  • Score as severity. Ranking a moderation queue by TOXICITY surfaces mild, unambiguous insults ahead of serious threats; route THREAT and SEVERE_TOXICITY separately.
  • Comparing numbers across scorers or dates. Leaderboards and internal dashboards that mix them show trends that are artefacts of the instrument.
  • Offset bugs. Treating UTF-16 span offsets as Python indices highlights the wrong text in any comment containing emoji.
  • Unsupported languages. Sending comments in a language an attribute does not support returns errors; a client that treats errors as zero scores publishes them unchecked.
  • Missing archive. Appeals and audits that need the original score cannot be answered after shutdown unless scores and versions were stored with the decision.

Trade-offs

Replacement routeStrengthCost
Open classifier you hostPinned weights, reproducible scores, no quotaServing, monitoring and bias testing are yours
Another hosted moderation APILittle to operate, broad category coverageVersion drift and another future deprecation
Safety model scoring with a policy promptUses your own written policy and contextHigher latency and cost per comment
Fine-tune on your moderator decisionsMatches your community's normsNeeds clean labels and retraining discipline

What to do next

  1. Find every call to comments:analyze and every threshold that consumes its output, and move thresholds into one policy configuration.
  2. Set doNotStore on all calls and make outage behaviour an explicit hold, not an exception path.
  3. Start storing raw scores with the scorer name and date next to every moderation decision.
  4. Before December 2026, rescore the samples, baselines and benchmark sets you will need later.
  5. Pick a replacement, run the replay with matched thresholds, and review disagreements by hand.
  6. Run the identity-term and dialect regression set against both scorers.
  7. Shadow the replacement on live traffic, cut over the policy, and delete the dead client code.
Key takeaway: Perspective scores estimate how many readers would find a comment toxic, under one model version, and the service ends on 31 December 2026. Keep thresholds in a policy layer, store raw scores with dates, rescore what you will need before the shutdown, and migrate by replaying real traffic and matching operating points against moderator decisions rather than by swapping the scorer.