If you ship a chatbot, an image generator or an LLM-powered feature to users in India, there is no single AI act to read. What applies instead is a stack of instruments: voluntary national AI governance guidelines, a binding amendment to the intermediary rules that covers synthetic media, a data-protection law whose core duties are still being phased in, and long-standing cyber-incident directions. Each has a different status, a different start date and a different enforcer.
This article maps that stack to engineering work. For each instrument it states what is in force as of October 2026, what is coming and when, and the control that satisfies it. It then works through a concrete product crossing the significant social media intermediary threshold, and lists failure modes and trade-offs. Dates and figures were checked against published legal and press summaries of the official texts in October 2026. This is engineering guidance, not legal advice; have counsel confirm how each rule applies to your product.
A stack of instruments, not one act
The Ministry of Electronics and Information Technology (MeitY) released the India AI Governance Guidelines on 5 November 2025 and chose explicitly not to propose a standalone AI law. The stated approach is to apply existing law (information technology, data protection, consumer protection and sectoral regulation) and to fill gaps with targeted amendments and voluntary measures. For a builder that means compliance work is spread across several texts, and the binding pieces are mostly about content and data, not about models as such.
The AI Governance Guidelines
The guidelines rest on seven principles, called sutras, adapted from the Reserve Bank of India's FREE-AI committee report: trust as the foundation, people first, innovation over restraint, fairness and equity, accountability, understandable by design, and safety, resilience and sustainability. They propose three institutions: an AI Governance Group to coordinate policy across ministries, a Technology and Policy Expert Committee to advise it, and the AI Safety Institute, already set up under the IndiaAI Mission, for testing and standards work.
None of this creates a direct legal duty, but it signals where binding rules will come from and what regulators will ask in an inquiry. The recurring asks are graded risk assessment, transparency reports, voluntary commitments, grievance redress and reporting of AI incidents to a national database. A team that already keeps a model card per deployed model, an incident register with severity and root cause, and a documented risk classification per use case can answer those asks without new work. A team that does not will find the binding instruments below harder too, because they need the same evidence.
Synthetic media rules, binding since February 2026
The binding change that matters most to generative AI products is the amendment to the Information Technology (Intermediary Guidelines and Digital Media Ethics Code) Rules, 2021, notified on 10 February 2026 and in force from 20 February 2026. It defines synthetically generated information (SGI) as audio, visual or audio-visual content that is artificially or algorithmically created, generated, modified or altered by a computer resource in a way that appears real and authentic. Good-faith editing such as formatting, colour adjustment, noise reduction and compression is excluded, as are accessibility, translation and discoverability uses that do not alter substance. Text-only chatbot output is not what this definition targets; generated or edited images, video, voice and music are.
Two groups carry duties. An intermediary that offers a tool for creating SGI must label permitted synthetic content prominently and embed permanent metadata or a provenance marker with a unique identifier that traces back to its computer resource, and must not let users modify, suppress or remove that marker. A significant social media intermediary, meaning a platform above the government-notified threshold of five million registered users in India, must also ask users to declare whether uploaded content is synthetic, deploy technical measures to verify those declarations, and label content confirmed as SGI.
The amendment also cut the takedown window. Unlawful content flagged by a court order or an authorised government notice must be removed or disabled within three hours, down from thirty-six, and complaints about non-consensual intimate imagery, including deepfake nudity, carry a two-hour window. The rules do not name a technical standard for the metadata. Content Credentials (C2PA) manifests are a reasonable implementation choice, but treat them as your choice, not as what the rule mandates.
Building the label, provenance and takedown pipeline
The engineering consequence is a label-and-provenance step on every generation path, plus a takedown queue with a clock. The sketch below shows the shape: every media output gets a visible label, an identifier recorded in your own store, and an embedded marker; a takedown order starts a timer that pages people before the deadline, not at it.
import hashlib, json, time, uuid
from dataclasses import dataclass
SGI_LABEL = "AI-generated" # visible label text, rendered on the media itself
@dataclass
class Provenance:
content_id: str # unique identifier traceable to our system
model: str
created_at: float
sha256: str
def finalize_media(raw: bytes, model: str, store) -> tuple[bytes, Provenance]:
prov = Provenance(str(uuid.uuid4()), model, time.time(),
hashlib.sha256(raw).hexdigest())
labelled = render_visible_label(raw, SGI_LABEL) # burn-in, not a tooltip
marked = embed_manifest(labelled, json.dumps(prov.__dict__)) # e.g. C2PA or watermark
store.put(prov.content_id, prov, sha256_out=hashlib.sha256(marked).hexdigest())
return marked, prov
TAKEDOWN_SLA = {"court_or_govt": 3 * 3600, "ncii": 2 * 3600}
def open_takedown(order, queue, pager):
deadline = order.received_at + TAKEDOWN_SLA[order.kind]
queue.add(order.id, deadline=deadline, targets=order.urls)
pager.schedule(at=deadline - 3600, msg=f"takedown {order.id}: 1h left")
pager.schedule(at=deadline - 900, msg=f"takedown {order.id}: 15m left", severity="high")Three details matter. Measure the deadline from when the order reaches you, so the intake channel (email alias, portal, legal inbox) must be monitored around the clock and must timestamp receipt. Removal must cover every copy and cache, including CDN edges and derived thumbnails, so keep a content-ID-to-object index. And log each step, because the proof that you met the window is the log.
DPDP: what starts when
The Digital Personal Data Protection Act, 2023 is India's general data-protection law, and the DPDP Rules notified on 13 November 2025 start it in phases. Rules 1, 2 and 17 to 21, which constitute the Data Protection Board, applied immediately. Rule 4 on consent managers applies one year after publication, around November 2026. Rules 3, 5 to 16, 22 and 23, which cover notice, security safeguards, breach intimation, retention, children's data and significant data fiduciaries, apply eighteen months after publication, around 13 May 2027. As of October 2026, the core duties are commencing, not yet binding, and that gap is your build window.
| Duty (from about May 2027) | What it means for an LLM product | Maximum penalty |
|---|---|---|
| Reasonable security safeguards | encryption, access control, logs, tested incident response | Rs 250 crore |
| Breach intimation | notify the Board and affected people without delay; detailed report within 72 hours | Rs 200 crore |
| Children (under 18) | verifiable parental consent; no tracking or targeted ads at children | Rs 200 crore |
| Significant data fiduciary duties | annual DPIA and audit; due diligence that algorithmic software does not risk users' rights | Rs 150 crore |
For LLM systems the hard parts are not the headline duties but their interaction with model pipelines. Prompts and chat histories are personal data when they identify someone, so consent notices must cover them, and retention must be bounded by purpose. Erasure requests must reach logs, evaluation sets and fine-tuning corpora, which needs lineage from each record to every derived dataset. The Act excludes personal data that the person themselves made publicly available, but that carve-out is narrower than most scraping pipelines assume, because data another person posted about someone is not covered by it unless that person was legally required to publish it; get a legal read before relying on it for training.
The control pattern is the same as under the EU AI Act and GDPR, which helps teams that serve both markets; compare the EU AI Act for LLM teams and data governance for LLMs.
CERT-In: six hours and 180 days
The CERT-In Directions issued on 28 April 2022 already apply. Covered entities, which include service providers, intermediaries and data centres, must report listed cyber incidents to CERT-In within six hours of noticing them, and keep logs of their ICT systems for a rolling 180 days within Indian jurisdiction. For an LLM service, prompt-injection-driven data exposure, account takeover and leaked API keys can all fall into reportable categories.
This collides with data minimisation: the logs you must keep contain the prompts you would rather discard. The usual resolution is to split logs into a security tier (request metadata, identities, tool calls, outcomes) retained for 180 days in an Indian region, and a content tier (full prompts and completions) with shorter, purpose-bound retention and redaction. Audit logging for LLM systems covers schema design for that split.
Worked example: a generative app with 6 million users
Consider a consumer app offering chat plus image and voice generation, with 6 million registered users in India, hosted partly in Mumbai and partly in Singapore. Walk it through the stack.
| Question | Answer for this product | Work item |
|---|---|---|
| Offers a tool that creates SGI? | yes: images and cloned voices | visible label plus embedded marker on every output |
| Above the SSMI threshold? | yes, 6 million is above 5 million, if users can share content with each other | upload declaration, verification classifier, labels |
| Takedown windows | 3 hours, 2 hours for intimate imagery | 24x7 intake, timed queue, CDN purge |
| CERT-In | yes, as a service provider | 180-day security logs in the Mumbai region; 6-hour runbook |
| DPDP, May 2027 | data fiduciary; possibly notified as significant | consent notices, erasure lineage, 72-hour breach runbook, DPIA |
| Children | minors likely present | age assurance and parental consent flow before May 2027 |
Sequenced by binding date, the order is: SGI labelling, upload declarations and the takedown clock first, because they apply today; CERT-In log placement next, because it is already in force and audits ask for it; then the DPDP programme, staged so consent and erasure plumbing ship before the May 2027 deadline rather than on it. The verification classifier deserves a measured error budget: publish its false-negative rate internally, because a declaration you never check will not satisfy a duty to verify.
Failure modes
- Labels that strip in transit. Re-encoding, resizing or screenshots remove embedded metadata. Pair metadata with a visible burn-in label and a server-side record keyed by content hash, so you can still prove origin.
- Takedown clock started at triage. If the order sat in an unmonitored inbox for two hours, you have one hour left. Timestamp at receipt.
- Erasure that misses derived data. Deleting the account row leaves prompts in evaluation sets and fine-tuning shards. Keep record-level lineage.
- Logs in the wrong region. Security logs replicated only to a foreign region fail the in-India retention expectation. Check where your logging vendor stores data.
- Treating voluntary as optional forever. The guidelines' incident and transparency asks are the template for future binding rules; skipping them now means rebuilding later.
- Assuming text is exempt from everything. Text output escapes the SGI definition, but chat logs are still personal data and incidents involving them are still reportable.
Trade-offs
One global pipeline or India-specific behaviour. Applying labels and provenance to every output worldwide is simpler and survives users crossing borders, at the cost of features some markets do not require. Geo-specific behaviour is cheaper per request but multiplies test surfaces. Most teams label globally and localise only takedown routing and data residency.
Metadata or watermark. Metadata manifests are rich and verifiable but fragile under re-encoding; pixel or audio watermarks survive transformations but carry little information and can be attacked. Using both, with a server-side ledger behind them, is the robust option.
Build now or wait for May 2027. Waiting saves effort if the rules change, but consent, lineage and erasure touch storage schemas, and those migrations take quarters. Track changes in one place, as described in LLM regulation in depth, and build the plumbing early.
What to do next
- List every output path that can produce images, video or audio, and confirm each applies a visible label and an embedded, server-recorded identifier.
- Check whether users can share content with each other on your service; if so, count registered users in India against the five-million threshold and, above it, build the declaration and verification step.
- Set up a monitored takedown intake that timestamps receipt, with timers for the three-hour and two-hour windows and a CDN purge path.
- Move security logs for Indian traffic to an Indian region with 180-day retention, and write a six-hour CERT-In reporting runbook.
- Start a DPDP programme with a May 2027 deadline: consent notices, erasure lineage across logs and training data, a 72-hour breach runbook, and age assurance.
- Keep a model card and incident register per deployed model so you can answer the governance guidelines' asks.
- Re-check MeitY notifications quarterly; this stack is still moving.