Every data platform eventually gets the same three questions. Where did this number come from? Which tables contain personal data? If I change this column, what breaks downstream? Apache Atlas is the Hadoop ecosystem's answer: a metadata store that records datasets, the processes that move data between them, and the classifications attached to both, then lets people and policy engines query that graph.

This article explains how Atlas works so you can run it, not just install it: the type system, how hooks feed metadata through Kafka, how lineage forms, how classifications propagate and how to stop them, the APIs you will script against, and the problems that appear after the first month. Version facts come from the ASF release catalog: 2.4.0 shipped in January 2025 and 2.5.0 in April 2026; 2.6.0 release candidates are tagged, but a final 2.6.0 release was not confirmed at the time of writing.

Advertisement

What Atlas is, and what it is not

Atlas is a catalog with a graph underneath. It stores entities (a Hive table, an HDFS path, a Kafka topic, a job run), relationships between them, and classifications such as PII or Finance that act as tags with attributes. It does not move or read your data, does not run queries, and does not measure data quality. It also does not enforce access control itself; it supplies tags that Apache Ranger can enforce through tag-based policies.

The value comes from two things a spreadsheet catalog cannot do: metadata that is captured automatically as jobs run, and classifications that follow data through lineage. Without hooks feeding it, Atlas is just a slow wiki, so most of this article is about keeping that feed healthy.

Architecture: hooks in, graph in the middle, events out

Apache Atlas: metadata in through hooks and REST, change events out through KafkaHive hookHiveServer2 / metastoreHBase, Kafka hooksSqoop, Storm, ImpalaCustom producersyour ETLKafkatopic ATLAS_HOOKAtlas serverType systemGraph engine (JanusGraph)Lineage, propagationREST v2 + UIconsumeHBasegraph storeSolror ElasticsearchKafkaATLAS_ENTITIESpublishRanger tagsyncYour consumersalerts, catalogUsers, scriptsREST, searchZooKeeperHA election for AtlasHooks are asynchronous by default: the producer keeps working when Atlas is slow, and Atlas catches up from Kafka
Producers publish metadata to ATLAS_HOOK; the Atlas server applies it to the graph and publishes change events, including classification changes, to ATLAS_ENTITIES.

There are two ways in. The REST API is synchronous: you send entities and get a mutation response. Hooks are the other path. A hook runs inside a producer such as HiveServer2, turns each operation into a notification and publishes it to the Kafka topic ATLAS_HOOK. The Atlas server consumes that topic and applies the changes. Hook message types are ENTITY_CREATE_V2, ENTITY_PARTIAL_UPDATE_V2, ENTITY_FULL_UPDATE_V2 and ENTITY_DELETE_V2, alongside older non-V2 forms.

Going out, Atlas publishes to ATLAS_ENTITIES whenever entities or classifications change: ENTITY_CREATE, ENTITY_UPDATE, ENTITY_DELETE, CLASSIFICATION_ADD, CLASSIFICATION_UPDATE and CLASSIFICATION_DELETE. Ranger's tagsync is the best-known consumer, but anything can subscribe, for example a job that alerts when a new table gains a PII tag.

Inside, the type system validates everything, the graph engine stores entities as vertices and edges in JanusGraph, classically backed by HBase, and keeps search indexes in Solr or Elasticsearch. Hooks exist in the Atlas source tree for Hive, HBase, Kafka, Sqoop, Storm and others; the 2.5 release notes also mention Impala, Spark and a Trino metadata extractor. Check the hook list for your exact version.

Advertisement

The type system: everything is a typed entity

Every object in Atlas is an instance of a type. Entity types describe things; classification types describe tags; relationship types describe typed edges; struct and enum types describe attribute shapes; business metadata adds curated attributes to existing types. Two built-in supertypes matter most: DataSet, for anything that holds data, and Process, which has inputs and outputs lists of datasets. Lineage is nothing more than Process entities connecting DataSets.

Uniqueness comes from qualifiedName. The Hive hook uses db@cluster for databases, db.table@cluster for tables and db.table.column@cluster for columns, all lowercase, where the cluster suffix comes from atlas.cluster.name. Hive entities are hive_db, hive_table, hive_column, hive_storagedesc, hive_process and hive_column_lineage. Follow the same convention in every integration: two tools that name one table differently create two entities and a hole in the lineage graph.

POST /api/atlas/v2/types/typedefs
{
  "entityDefs": [
    { "name": "etl_job_run", "superTypes": ["Process"],
      "attributeDefs": [
        { "name": "gitCommit",  "typeName": "string", "isOptional": true },
        { "name": "rowsWritten", "typeName": "long",  "isOptional": true }
      ] }
  ],
  "classificationDefs": [
    { "name": "PII", "description": "Contains personal data",
      "attributeDefs": [ { "name": "level", "typeName": "string", "isOptional": true } ] }
  ]
}

Model your own jobs as subtypes of Process and give them attributes you will actually query, such as the code version. Keep the process qualifiedName stable per job, not per run, or the graph fills with one vertex per execution.

Hooks and the notification path

The Hive hook is a post-execution hook. After a statement succeeds, it inspects the query plan's inputs and outputs, builds entities for touched databases, tables and columns, and for statements that move data, a hive_process with column lineage. Column lineage records a dependency type of SIMPLE, EXPRESSION or SCRIPT, describing how each output column derives from its inputs.

<!-- hive-site.xml on HiveServer2 -->
<property>
  <name>hive.exec.post.hooks</name>
  <value>org.apache.atlas.hive.hook.HiveHook</value>
</property>

# atlas-application.properties next to the hook
atlas.cluster.name=prod1                 # becomes the @prod1 suffix of every qualifiedName
atlas.hook.hive.synchronous=false        # recommended: send from a background queue
atlas.hook.hive.numRetries=3
atlas.hook.hive.queueSize=10000
atlas.kafka.bootstrap.servers=kafka1:9092,kafka2:9092

# one-off backfill of tables that existed before the hook was installed
<atlas hook package>/hook-bin/import-hive.sh

With atlas.hook.hive.synchronous=false, the recommended setting, the hook queues notifications and a background thread sends them, so a slow Kafka or Atlas never stalls a query. The cost is that metadata is eventually consistent: a table created a second ago may not be searchable yet. The queue is bounded by queueSize, and sends are retried numRetries times. A notification that never reaches Kafka never reaches Atlas, so watch the hook's log on the producer hosts as well as Atlas itself.

Hooks only see operations that happen after they are installed. Use import-hive.sh to bootstrap existing databases, filtering with database and table patterns or a file, and rerun it after any period in which the hook was broken. On a Kerberized cluster the hook and the Atlas server need service principals and the Kafka topics need ACLs; Hadoop Kerberos covers the principal setup.

Lineage: Process entities connecting DataSets

Atlas builds lineage from the inputs and outputs of Process entities. A CREATE TABLE AS SELECT in Hive becomes a hive_process whose inputs are the source tables and whose output is the new one. For anything without a hook, such as a Python ETL job, you publish the process yourself. The example below uses only REST endpoints present in Atlas's EntityREST and LineageREST classes, referencing existing tables by unique attribute rather than looking up GUIDs.

import requests

ATLAS = "https://atlas.example.com:21443/api/atlas/v2"
S = requests.Session()
S.auth = ("svc_etl", "...")          # or SPNEGO on a Kerberos cluster

def ref(type_name, qualified_name):
    """Reference an existing entity by its unique attribute instead of its GUID."""
    return {"typeName": type_name, "uniqueAttributes": {"qualifiedName": qualified_name}}

run = {
    "typeName": "etl_job_run",
    "attributes": {
        "qualifiedName": "nightly_orders_to_mart@prod1",   # stable: one process per job, not per run
        "name": "nightly_orders_to_mart",
        "inputs":  [ref("hive_table", "raw.orders@prod1"), ref("hive_table", "raw.customers@prod1")],
        "outputs": [ref("hive_table", "mart.daily_orders@prod1")],
        "gitCommit": "4f2c9e1",
        "rowsWritten": 198800,
    },
}
r = S.post(f"{ATLAS}/entity", json={"entity": run}, timeout=30)
r.raise_for_status()
print(r.json().get("mutatedEntities", {}))

# read it back: two hops upstream of the mart table
mart = S.get(f"{ATLAS}/entity/uniqueAttribute/type/hive_table",
             params={"attr:qualifiedName": "mart.daily_orders@prod1"}, timeout=30).json()
guid = mart["entity"]["guid"]
lineage = S.get(f"{ATLAS}/lineage/{guid}", params={"direction": "INPUT", "depth": 2}, timeout=30).json()
for edge in lineage["relations"]:
    print(edge["fromEntityId"], "->", edge["toEntityId"])

The lineage endpoint takes direction (INPUT, OUTPUT or BOTH) and depth; the defaults in the source are BOTH and 3. Deep, wide graphs make lineage queries expensive, so ask for the direction you need with a small depth and expand interactively.

Classification propagation, and how to contain it

Classification propagation is Atlas's most useful and most surprising feature. When an entity carries a classification with propagation enabled, Atlas attaches it to every entity downstream through lineage. Tag a raw table PII, and the views and CTAS tables derived from it inherit PII automatically, which is how tags keep up with data that analysts copy around. Updates and removals propagate too, and deleting an entity in the middle of a path removes propagated tags downstream unless another path still carries them.

Propagation is on by default for each new association. There are three controls: the propagate flag on the individual classification association; a propagate flag on each lineage edge, also on by default; and a list of blocked propagated classifications on an edge. The last is what you use when a view masks a PII column: block PII on that edge, and the view stops inheriting it.

  • Over-propagation: one tag on a widely joined dimension table can flow to hundreds of downstream tables, and Ranger may then mask or deny access to all of them. Tag at the column level where possible, and review propagation reach before tagging hubs.
  • Event fan-out: Atlas publishes an ATLAS_ENTITIES notification for each entity a propagation touches, so one classification change can produce thousands of events for tagsync and other consumers.
  • Silent removal: dropping an intermediate table removes propagated tags downstream if no other path exists. Treat lineage deletions as security-relevant.

Search, glossary and business metadata

Atlas offers basic search, with filters on type, classification and attributes; a DSL for more structured queries; full-text search; and relationship search. The glossary maps business terms such as "active customer" onto technical entities, and business metadata attaches curated attributes, such as a data owner or retention class, to existing types without redefining them.

# basic search: every hive_table carrying PII, without deleted entities
GET /api/atlas/v2/search/basic?typeName=hive_table&classification=PII&excludeDeletedEntities=true&limit=100

# DSL search: tables in one database
GET /api/atlas/v2/search/dsl?query=hive_table where qualifiedName like 'mart.*'

# tag a column explicitly, with propagation switched off for this association
POST /api/atlas/v2/entity/guid/{guid}/classifications
[ { "typeName": "PII", "attributes": { "level": "direct" }, "propagate": false } ]

# which instance is serving?  ACTIVE on the leader, PASSIVE on the others
GET /api/atlas/admin/status

Basic search with a classification filter is the quickest way to answer "which tables are tagged PII?" for an auditor.

Running it: HA, storage and the metrics that matter

Atlas runs active/passive. Several instances register with ZooKeeper (see ZooKeeper in Hadoop), one is elected active, and passive instances redirect user requests to the active one with an HTTP redirect. /api/atlas/admin/status reports ACTIVE, PASSIVE or a transition state, and the Atlas documentation shows an HAProxy health check that routes only to the instance reporting ACTIVE. Because only one instance serves, HA protects availability but does not add throughput. Size the active node for peak ingest.

Watch three things. First, consumer lag on ATLAS_HOOK: rising lag means metadata is falling behind reality, and it is the earliest sign of an overloaded Atlas. Second, hook errors on producer hosts, because those messages never arrive. Third, growth of soft-deleted entities: by default, deletes mark entities as deleted rather than removing them, and they accumulate in the graph and index. The 2.5 release added an automatic purge capability; on older versions plan a manual purge procedure. Back up HBase and Solr together, since a graph restored without its matching index returns search results that do not match entities.

Worked example: a PII tag from raw to report

A team lands raw.customers daily. A steward adds PII to the email and phone columns, with propagation on. An analyst runs CREATE TABLE mart.customer_orders AS SELECT c.email, o.total FROM raw.customers c JOIN raw.orders o .... The Hive hook publishes a hive_process with column lineage; Atlas links the new email column to the source column and propagates PII to it once it consumes the message. Atlas emits CLASSIFICATION_ADD to ATLAS_ENTITIES, Ranger tagsync picks it up, and a tag-based masking policy for PII now applies to mart.customer_orders.email without anyone writing a resource policy for the new table.

Later, a reporting view hashes the email. The steward adds PII to the blocked list on that lineage edge, and the hashed column stops carrying the tag. That sequence makes a good acceptance test for a new deployment.

Failure modes and trade-offs

  • Hook installed on one HiveServer2 but not another: lineage has holes that look like data appearing from nowhere. Configure hooks with the same automation as the service.
  • Inconsistent qualifiedName: a custom producer that omits the cluster suffix or keeps uppercase creates duplicate entities.
  • Kafka outage longer than the hook queue: notifications are lost; rerun the import utility for the affected window.
  • Graph and index drift after a partial restore: search finds entities that fail to load, or misses existing ones.
  • Per-run process entities: lineage queries slow down as the graph grows by one vertex per job execution.

The trade-off is depth against breadth. Atlas has deep Hadoop integration, column lineage from Hive, and tag-based security through Ranger that few catalogs match on-premises. It is also an operationally heavy system with its own HBase, Solr, Kafka and ZooKeeper dependencies, and its hooks are strongest for the classic Hadoop stack. If your platform has moved to a lakehouse on object storage, compare it with the catalog your table format and engines already use, as discussed in the modern lakehouse, before committing to it.

What to do next

  1. Pick an atlas.cluster.name and a lowercase qualifiedName convention, write them down, and apply them to every producer.
  2. Install the Hive hook on every HiveServer2 with synchronous=false, then backfill with import-hive.sh.
  3. Model one non-Hive ETL job as a Process subtype and publish its lineage through the REST API.
  4. Alert on ATLAS_HOOK consumer lag and on hook errors in producer logs.
  5. Define classification types, decide where propagation should stop, and record blocked edges for masked views.
  6. Connect Ranger tagsync and test one tag-based policy end to end with the worked example above, following Hive on Hadoop for the table side.
  7. Put both Atlas instances behind a health check on /api/atlas/admin/status, and schedule soft-delete purges and joint HBase and Solr backups.
Key takeaway: Apache Atlas stores datasets, processes and classifications as a typed graph. Hooks publish metadata asynchronously to the ATLAS_HOOK Kafka topic, Atlas applies it to JanusGraph with a Solr or Elasticsearch index, and change events go out on ATLAS_ENTITIES to consumers such as Ranger tagsync. Lineage is Process entities linking DataSets, and classifications propagate along it unless you turn propagation off or block it per edge. Keep qualifiedName conventions consistent, backfill when hooks fail, watch hook consumer lag, and remember that HA is active/passive.