PIPEDA, the Personal Information Protection and Electronic Documents Act, is Canada's federal private-sector privacy law. It predates generative AI by two decades and contains no AI-specific rules, yet it governs how a Canadian business may collect, use and disclose personal information through an LLM product: the prompts users type, the documents a retrieval system indexes, the logs you keep and the data you might fine-tune on.

This is no longer theoretical. On 6 May 2026 the federal Privacy Commissioner and the Quebec, British Columbia and Alberta regulators released joint findings on OpenAI's ChatGPT. They concluded that its initial training did not comply with Canadian privacy law, mainly on consent and transparency. This article explains PIPEDA from first principles, maps its ten principles onto an LLM architecture, walks through a worked deployment with code, and covers breach reporting, the provincial overlay, and where federal reform stands. It is engineering guidance, not legal advice; confirm your position with counsel.

Who and what PIPEDA covers

PIPEDA applies to organisations that collect, use or disclose personal information in the course of commercial activity, and to federally regulated works such as banks, telecoms and airlines, including their employee data. Personal information is information about an identifiable individual, a test the regulator reads broadly: a prompt containing a customer's account history is personal information, and so is a model output that makes claims about a named person.

Quebec, Alberta and British Columbia have their own private-sector laws that the federal government has declared substantially similar. Inside those provinces, the provincial law generally governs intra-provincial activity, while PIPEDA still applies to federally regulated businesses and to personal information that crosses provincial or national borders. An LLM service with users across Canada therefore usually answers to PIPEDA and to one or more provincial laws at once, which is exactly why the ChatGPT investigation was joint.

The core rule is section 5(3): an organisation may collect, use or disclose personal information only for purposes that a reasonable person would consider appropriate in the circumstances. Consent does not cure an inappropriate purpose. In the ChatGPT findings, the regulators accepted that building a large language model can be an appropriate purpose; the failures were in how the information was obtained and explained.

The ten principles as engineering requirements

Schedule 1 of PIPEDA sets out ten fair information principles. Each one becomes a concrete engineering requirement once you draw the data flows of an LLM system.

PrincipleWhat it means for an LLM product
1. AccountabilityName a privacy officer; you stay responsible for data you send to a model vendor, so contracts must bind them
2. Identifying purposesState why prompts, logs and documents are processed before or at collection
3. ConsentMeaningful consent for each purpose; training on conversations is a separate purpose from answering them
4. Limiting collectionCollect only what the feature needs; redact identifiers before prompts leave your boundary
5. Limiting use, disclosure, retentionNo reuse of logs for new purposes without consent; delete on a schedule
6. AccuracyOutputs about people can be wrong; warn users and provide a correction path
7. SafeguardsAccess control, encryption, prompt-injection defences proportional to sensitivity
8. OpennessPublish plain-language policies describing the model, vendors and data flows
9. Individual accessUsers can request what you hold about them, including logs and stored memories
10. Challenging complianceA complaint process that reaches someone empowered to act

Consent, purposes and scraped data

Since 2018 the Commissioner's guidelines on meaningful consent have required organisations to make the key elements prominent: what is collected, for which purposes, with whom it is shared, and what the risks of harm are. Section 6.1 makes consent valid only if the individual would reasonably understand the nature, purpose and consequences. Express consent is generally needed for sensitive information such as health and financial data, or for uses outside reasonable expectations.

Two LLM-specific consequences follow. First, scraped web data is not a free pass. PIPEDA's exception for publicly available information is defined narrowly by regulation: telephone directories, professional listings, registries, and publications such as books and newspapers where the individual provided the information. Social media and forum posts are not on that list, which is the reasoning the regulators applied to Clearview AI in 2021 and to ChatGPT in 2026. Second, using customer conversations to improve a model is a new purpose. Bundling it into a terms-of-service click is the pattern the ChatGPT findings criticised.

The 2026 ChatGPT findings

The joint investigation, opened in 2023, examined how OpenAI collected personal information from publicly accessible internet sources, licensed datasets and user interactions to train GPT-3.5 and GPT-4. The regulators found deficiencies in consent, transparency, overcollection of sensitive information, accuracy notices, access and deletion mechanisms, and accountability, noting that the product was released without known privacy risks being fully addressed.

Outcomes differed by regulator, which is instructive in itself. The federal Commissioner found the complaint well-founded and conditionally resolved, relying on OpenAI's commitments: filtering to mask personal information in training data, measures to block specific personal details about public figures, improved handling of access requests, user notices and periodic compliance reporting. British Columbia and Alberta found the complaints well-founded and unresolved, British Columbia holding that models trained on scraped data contravene its consent rules, which differ from PIPEDA's. Quebec's findings were mixed, unresolved on consent. The lesson for builders: a remedy that satisfies the federal regulator may not satisfy a province.

Architecture and data flow

A PIPEDA-aware LLM product: every arrow is a purpose, a consent basis and a retention clockCanadian usernotice, consent choiceGatewayPII redaction, purpose tagLLM providerprocessor, maybe abroadResponseaccuracy noticepromptminimisedRetrieval storepurpose-filtered docsConversation logretention clockTraining setseparate consentcontextlogopt-in onlyAccess and correction30-day responseBreach registerRROSH test, 24-month recordsAccountabilityprivacy officer, vendor contractsfind and deletePIPEDA has no AI chapter. The ten principles apply to each box and arrow above.
Data flows of a compliant assistant. Redaction happens before the vendor boundary, training is a separate opt-in sink, and access, breach and accountability processes span every store.

Read the diagram as a list of obligations. Every store needs a retention period and must be searchable by user for access requests. Every arrow that leaves Canada needs a contract and a notice. PIPEDA does not prohibit cross-border transfers; it treats them as a use by a service provider, keeps you accountable, and expects you to tell users their data may be processed in another jurisdiction and be accessible to its authorities.

Worked example: a bank's retrieval assistant

A federally regulated Canadian bank deploys an assistant that answers customer questions using the customer's own account documents through retrieval. The design decisions, principle by principle: the purpose is answering that customer's questions; retrieval must never surface another customer's documents; identifiers are tokenised before prompts reach a US-hosted model; conversations are kept 90 days for dispute handling and then deleted; and nothing is used for training unless the customer opts in separately. The gateway enforces purpose and minimisation in code rather than in policy documents:

import re, hashlib, time

SIN = re.compile(r"\b\d{3}[- ]?\d{3}[- ]?\d{3}\b")      # Social Insurance Number shape
CARD = re.compile(r"\b(?:\d[ -]?){13,19}\b")
RETENTION_S = 90 * 86400

def tokenise(text, vault, user_id):
    def sub(m):
        tok = "TOK_" + hashlib.sha256((user_id + m.group()).encode()).hexdigest()[:10]
        vault[tok] = m.group()              # stays in Canada, never sent to the vendor
        return tok
    return CARD.sub(sub, SIN.sub(sub, text))

def retrieve(store, user_id, query, purpose="answer_own_account"):
    # Purpose binding: the filter is mandatory, not a ranking hint.
    return store.search(query, filter={"owner": user_id, "purposes": purpose}, k=5)

def handle(user, prompt, store, llm, log, vault):
    clean = tokenise(prompt, vault, user.id)
    docs = [tokenise(d.text, vault, user.id) for d in retrieve(store, user.id, clean)]
    answer = llm.complete(system=POLICY, context=docs, prompt=clean)
    log.write({"user": user.id, "prompt": clean, "answer": answer,
               "expires": time.time() + RETENTION_S,
               "train_ok": user.consents.get("model_improvement", False)})
    return detokenise(answer, vault)

The log record carries its own expiry and consent flag, so the training-export job can filter on train_ok and the deletion job on expires without consulting another system. An access request becomes a query by user across the log, the vault and any long-term memory store; PIPEDA requires a response within 30 days, extendable in limited cases. The tokenising regexes are deliberately crude and will miss free-text identifiers such as names and addresses, so production systems add an entity-recognition pass and measure its recall on labelled samples.

Breaches and enforcement

Since November 2018, PIPEDA requires organisations to report to the Commissioner, and notify affected individuals, any breach of security safeguards that creates a real risk of significant harm (RROSH). The assessment weighs the sensitivity of the information and the probability of misuse. Every breach, reportable or not, must be recorded and the records kept for 24 months.

LLM systems add new breach paths: a prompt injection that makes the assistant reveal another customer's retrieved document, a logging pipeline that copies unredacted prompts into an analytics tool, or a vendor incident exposing stored conversations. Each needs to be in your incident playbook with an RROSH template.

Here PIPEDA differs sharply from GDPR. It has no general administrative penalty regime. The Commissioner investigates and recommends; enforcement runs through compliance agreements, the Federal Court, and a small set of offences, such as knowingly failing to report or record a breach, carrying fines of up to $100,000. The reputational cost of a published finding is the main lever, but Quebec is a different story.

The provincial overlay: Quebec, Alberta, British Columbia

Quebec's Law 25 amended the province's private-sector act in stages between 2022 and 2024 and is the strictest regime in Canada. Three provisions matter most for LLM products. When a decision about a person is based exclusively on automated processing, the organisation must inform the person, and on request explain the personal information and the principal factors used, with a way to have the decision reviewed by a human. Before communicating personal information outside Quebec, the organisation must carry out a privacy impact assessment, which applies to many model APIs hosted elsewhere. And technology that can identify, locate or profile a person must be off by default. Penalties reach $10 million or 2 percent of worldwide turnover administratively, and up to $25 million or 4 percent for penal offences.

Alberta's and British Columbia's PIPA laws have their own consent rules, and as the ChatGPT findings showed, they can reach different conclusions on the same facts. Design for the strictest regime you operate in.

Federal reform: where it stands

Bill C-27, which would have replaced PIPEDA with the Consumer Privacy Protection Act and added an Artificial Intelligence and Data Act, died when Parliament was prorogued in January 2025. In June 2026 the government tabled Bill C-36, which would replace PIPEDA with a retitled privacy statute and restructure oversight; AI regulation is not part of it. At the time of writing the bill is before Parliament and not in force. PIPEDA remains the law. Build to its principles now and track the bill's text, but do not design against provisions that may change in committee.

Failure modes

FailureWhy it breaches PIPEDAFix
Training on chats by defaultNew purpose without meaningful consentSeparate, revocable opt-in; filter exports on it
Fine-tuning on scraped profilesOutside the narrow publicly available exceptionLicensed or consented sources; masking filters
Logs kept foreverRetention beyond the identified purposeExpiry stamped on every record; deletion job with evidence
Cross-tenant retrievalDisclosure without consent; a reportable breachMandatory owner filter; tests that try to cross it
No way to find a user's dataAccess requests cannot be answered in 30 daysIndex every store by user identifier
Model invents facts about a personAccuracy principleAccuracy notices, correction path, suppression list
Silent vendor in another countryOpenness and accountabilityContracts, notices, and for Quebec a transfer PIA

Trade-offs

Redaction versus quality. Tokenising identifiers protects users but can hurt answers that depend on the redacted values; keep the mapping local and detokenise only in the final response.

Canadian hosting versus model choice. In-country hosting simplifies notices and Quebec assessments but may limit which models you can use. Neither PIPEDA nor Law 25 mandates residency, so this is a risk and procurement decision.

Opt-in training versus data volume. Opt-in yields less data, but data obtained without valid consent is a liability that may need to be purged from models later, which is far more expensive.

Logging for safety versus minimisation. Abuse monitoring needs logs; keep them short-lived, redacted and separate from analytics.

What to do next

  1. Draw your LLM data flows like the diagram above and write the purpose, consent basis and retention period on every arrow.
  2. Separate the consent for using conversations to train or evaluate models from the consent to provide the service.
  3. Audit training and fine-tuning sources against the publicly available information regulation; drop what does not fit.
  4. Put a mandatory owner filter on retrieval and write tests that try to read another user's documents.
  5. Make every store searchable by user so you can answer an access request within 30 days.
  6. Add LLM-specific scenarios to your breach playbook, with an RROSH assessment template and a 24-month register.
  7. If you serve Quebec, run a transfer PIA for each foreign model vendor and add automated-decision notices.
  8. Track Bill C-36 and read the full May 2026 ChatGPT findings on the Commissioner's site.
  9. Related reading: GDPR and LLMs, CCPA and LLMs, PII handling in LLM systems, PII leakage and the EU AI Act.
Key takeaway: PIPEDA has no AI chapter, but its ten principles reach every prompt, log and training set. Use personal information only for purposes a reasonable person would accept, get separate meaningful consent for training, treat scraped data as outside the publicly available exception, enforce purpose and retention in code, and design for Quebec's stricter rules. The 2026 ChatGPT findings show regulators will apply these old principles to new models.