VPC Flow Logs answer the first question in almost every AWS network investigation: did the packets get there, and were they allowed? Each record summarizes the traffic one network interface saw for one 5-tuple (source address, destination address, source port, destination port, protocol) during a short window, along with byte and packet counts and whether security groups and network ACLs accepted or rejected it. There is no payload, no HTTP path and no DNS name, just enough metadata to reconstruct who talked to whom, how much, and what got blocked.

This article explains the record format field by field, including the easy-to-misread ones. It covers what is never logged, how to choose between CloudWatch Logs, S3 and Firehose, and how to build an Athena table over S3 that stays cheap as data grows. It ends with queries for the investigations teams actually run: rejected traffic, top talkers, internet egress and the classic security group versus network ACL puzzle.

How flow logs are captured

Where a flow record comes from and where it goesENI (capture point)per interface, per 5-tupleFlow log on VPC,subnet or ENIAggregation window10 min or 1 min maxpacketsCloudWatch Logs~5 min, InsightsAmazon S3~10 min, gzip filesData Firehosestream to SIEMAthenaSQL, projectionNot recorded at allAmazon DNS, IMDS, time sync, DHCP, ARP ...Fixed at creationformat, destination, interval, IAM roleA record is metadata about packets, never payload. The format cannot be edited later: delete and recreate.
Flow logs are captured at each elastic network interface, aggregated per window, and delivered to one destination per subscription.

You attach a flow log to a VPC, a subnet or a single network interface (ENI). Whatever you pick, capture happens per ENI: a VPC-level flow log is a convenient way of saying every ENI in this VPC, including ones created later by load balancers, NAT gateways, Lambda functions and VPC endpoints. Each subscription chooses a traffic filter (ACCEPT, REJECT or ALL), one destination and a maximum aggregation interval of 10 minutes (the default) or 1 minute. On Nitro-based instances the interval is always 1 minute or less regardless of the setting.

Delivery is best effort. AWS documents typical delivery of about 5 minutes to CloudWatch Logs and about 10 minutes to S3, on top of the aggregation window, so flow logs are a forensic and analytics tool, not a real-time intrusion prevention feed. Two more constraints shape every design decision: you cannot change a flow log's format, destination or IAM role after creation (you delete it and create a new one), and you can only enable flow logs on peered VPCs that live in your own account.

Anatomy of a flow record

The default format contains the 14 version 2 fields in a fixed order. A real record looks like this:

2 123456789012 eni-0a1b2c3d4e5f60718 10.0.1.25 10.0.2.40 49152 5432 6 12 3100 1728390000 1728390060 ACCEPT OK

Read left to right: format version 2, the account that owns the ENI, the ENI, source address, destination address, source port 49152, destination port 5432 (PostgreSQL), protocol 6 (TCP, by IANA number; 17 is UDP and 1 is ICMP), 12 packets, 3,100 bytes, window start and end in Unix seconds, the action and the log status.

FieldVersionWhat to watch for
action2ACCEPT or REJECT. REJECT covers security groups, network ACLs, and packets arriving after the connection closed.
log-status2OK, NODATA (no traffic in the window) or SKIPDATA (records dropped by an internal capacity limit or error). Count SKIPDATA before trusting totals.
start / end2First and last packet seen within the window, not the connection lifetime. Long connections produce one record per window.
tcp-flags3Bitmask of FIN 1, SYN 2, RST 4, SYN-ACK 18, OR-ed across the window. ACK and PSH are not recorded, so steady data transfer shows 0.
pkt-srcaddr / pkt-dstaddr3Original packet addresses. They differ from srcaddr/dstaddr behind a NAT gateway or for EKS pods using secondary IPs.
flow-direction5ingress or egress relative to the capturing ENI.
traffic-path5Egress path: 1 same VPC, 2 internet gateway or gateway endpoint, 3 virtual private gateway, 4 and 5 intra- and inter-region peering, 6 Local or Wavelength Zone, 7 gateway endpoint, 8 internet gateway.
reject-reason8Only BPA (VPC Block Public Access) or EC (encryption controls); a dash for every other reason, including security groups.

A custom format lets you choose any subset of fields in any order; its version is the highest version among the fields you chose. Missing or inapplicable values print as a dash. Two consequences catch people out. First, one TCP connection appears at least twice per window, once per direction, and if both ends are in your VPC each ENI logs both directions, so naive byte sums double count. Second, srcaddr for outbound traffic is always the ENI's primary private address; if you need the address the packet actually carried, you must have included pkt-srcaddr when you created the flow log, because you cannot add it afterwards.

What flow logs never see

Flow logs do not capture all IP traffic. The documented exclusions are traffic to the Amazon-provided DNS server (traffic to your own DNS server is logged), Windows license activation, the instance metadata service at 169.254.169.254, Amazon Time Sync at 169.254.169.123, DHCP, ARP, the source side of traffic mirroring, traffic to the VPC router's reserved address, and traffic between an endpoint ENI and a Network Load Balancer ENI. The DNS gap matters for security work: a host exfiltrating data through DNS queries to the Amazon resolver leaves no flow records, so pair flow logs with Route 53 Resolver query logs or GuardDuty, which consumes both sources.

Choosing a destination and a format

DestinationLatencyBest forWatch out for
CloudWatch Logs~5 minShort-retention alerting, metric filters, ad hoc Logs Insights queriesIngestion cost at high volume; set a retention policy or logs live forever
Amazon S3~10 minLong retention, Athena, cheapest per GB, lifecycle to colder storageBucket policy must allow the log delivery service; query design decides cost
Amazon Data FirehosestreamingFeeding a SIEM, OpenSearch or a third-party analytics toolYou own buffering, transformation and failed-delivery handling

Most organisations end up with S3 as the system of record, often a central logging account bucket that every workload account delivers to, plus a narrow CloudWatch Logs flow log (REJECT only, sensitive subnets only) for alarms. S3 delivery supports plain text or Parquet files, Hive-compatible prefixes and hourly partitions. AWS states that Parquet queries run 10 to 100 times faster than text and take about 20 percent less space, with the caveat that tiny volumes (under about 100 KB per period) can be larger in Parquet. Because the format is immutable, settle the field list before the first deployment:

aws ec2 create-flow-logs \
  --resource-type VPC --resource-ids vpc-0abc1234def567890 \
  --traffic-type ALL \
  --max-aggregation-interval 60 \
  --log-destination-type s3 \
  --log-destination arn:aws:s3:::acme-flow-logs/prod \
  --log-format '${version} ${account-id} ${interface-id} ${srcaddr} ${dstaddr} ${srcport} ${dstport} ${protocol} ${packets} ${bytes} ${start} ${end} ${action} ${log-status} ${vpc-id} ${subnet-id} ${instance-id} ${tcp-flags} ${pkt-srcaddr} ${pkt-dstaddr} ${flow-direction} ${traffic-path}'

This format keeps the 14 default fields in their default order and appends eight more, so tools that parse version 2 positionally still read the first 14 columns correctly. Single quotes stop the shell from expanding the ${...} placeholders.

An Athena table that stays cheap

With the default S3 layout, files land under prefix/AWSLogs/<account>/vpcflowlogs/<region>/<yyyy>/<mm>/<dd>/ as gzip text with a header line. The expensive mistake is a table without partitions, which scans the whole history for every query. The better pattern is partition projection: Athena computes partition locations from table properties, so there is no MSCK REPAIR TABLE or Glue crawler to run and new days are queryable as soon as files arrive. Columns must match the custom format order exactly, because text records are parsed by position:

CREATE EXTERNAL TABLE vpc_flow_logs (
  version int, account_id string, interface_id string,
  srcaddr string, dstaddr string, srcport int, dstport int,
  protocol bigint, packets bigint, bytes bigint,
  `start` bigint, `end` bigint, action string, log_status string,
  vpc_id string, subnet_id string, instance_id string, tcp_flags int,
  pkt_srcaddr string, pkt_dstaddr string,
  flow_direction string, traffic_path int
)
PARTITIONED BY (region string, day string)
ROW FORMAT DELIMITED FIELDS TERMINATED BY ' '
LOCATION 's3://acme-flow-logs/prod/AWSLogs/123456789012/vpcflowlogs/'
TBLPROPERTIES (
  'skip.header.line.count' = '1',
  'projection.enabled' = 'true',
  'projection.region.type' = 'enum',
  'projection.region.values' = 'us-east-1,eu-west-1',
  'projection.day.type' = 'date',
  'projection.day.format' = 'yyyy/MM/dd',
  'projection.day.range' = '2026/01/01,NOW',
  'projection.day.interval' = '1',
  'projection.day.interval.unit' = 'DAYS',
  'storage.location.template' =
    's3://acme-flow-logs/prod/AWSLogs/123456789012/vpcflowlogs/${region}/${day}'
);

Dashes in numeric columns become NULL, which is what you want for ICMP records with no ports. Every query should filter on day (and region if you have several), otherwise projection cannot prune and you pay for a full scan. If you chose Parquet, take the column definitions from the Athena integration that the VPC console can generate rather than guessing names.

Queries for real investigations

Four queries cover most investigations. Rejected traffic by destination port shows what is being blocked and whether it looks like scanning or a broken dependency:

SELECT dstport, protocol, count(*) AS records, count(DISTINCT srcaddr) AS sources
FROM vpc_flow_logs
WHERE day = '2026/10/08' AND action = 'REJECT' AND flow_direction = 'ingress'
GROUP BY dstport, protocol
ORDER BY records DESC
LIMIT 20;

Many distinct sources on 22, 3389 or 445 from public addresses is internet background noise that security groups are already stopping. A handful of internal sources on one port usually means a real service is misconfigured. Top talkers by bytes, counting egress only so each flow is counted once per ENI:

SELECT srcaddr, dstaddr, dstport, sum(bytes) / 1e9 AS gb
FROM vpc_flow_logs
WHERE day BETWEEN '2026/10/01' AND '2026/10/08' AND flow_direction = 'egress'
GROUP BY srcaddr, dstaddr, dstport
ORDER BY gb DESC
LIMIT 25;

Internet egress, which drives data transfer bills and is the first place to look for exfiltration, uses traffic_path = 8 (internet gateway). For instances behind a NAT gateway the instance's own record shows the egress path to the NAT ENI, so query the NAT gateway's ENI and group by pkt_srcaddr to see which private hosts are behind the volume:

SELECT pkt_srcaddr AS private_host, pkt_dstaddr AS internet_dest, sum(bytes) / 1e9 AS gb
FROM vpc_flow_logs
WHERE day = '2026/10/08' AND interface_id = 'eni-0nat1234567890abc'
  AND flow_direction = 'ingress'
GROUP BY pkt_srcaddr, pkt_dstaddr
ORDER BY gb DESC
LIMIT 25;

Worked example: security group or network ACL?

An application in subnet A reports intermittent timeouts connecting to PostgreSQL at 10.0.2.40 in subnet B. The security groups look right: the database group allows 5432 from the application group. Query both ENIs for that pair:

SELECT interface_id, srcaddr, srcport, dstaddr, dstport, action, flow_direction, tcp_flags
FROM vpc_flow_logs
WHERE day = '2026/10/08'
  AND ((srcaddr = '10.0.1.25' AND dstaddr = '10.0.2.40')
    OR (srcaddr = '10.0.2.40' AND dstaddr = '10.0.1.25'))
ORDER BY "start";

The output shows the inbound SYN to 5432 as ACCEPT on the database ENI, but the reply from port 5432 to the client's ephemeral port as REJECT on egress. Security groups are stateful: if they allowed the request, the reply is allowed automatically. Network ACLs are stateless and evaluate the reply as an independent packet. So the rejection must come from subnet B's outbound ACL, and the intermittency gives away the cause: someone allowed outbound 1024-49151 instead of the full ephemeral range, and clients only fail when Linux picks a source port above 49151 (its default range runs to 60999). The fix is an outbound ACL rule for 1024-65535 to subnet A. The same reasoning works in reverse: an ingress REJECT with no matching earlier ACCEPT is a security group or inbound ACL, and you check the ACL first because it is the only one of the two that can reject reply traffic.

Failure modes

  • Bucket policy rejects delivery. S3 delivery needs the log delivery service principal allowed to write to the prefix. In cross-account setups the flow log shows an error status in the console while the bucket stays empty; check it the day you create the subscription, not during an incident.
  • Missing fields discovered too late. No pkt-srcaddr means NAT and EKS traffic cannot be attributed. Because the format is immutable, the fix is a new flow log, and history before that day stays incomplete.
  • Double-counted bytes. Summing all records counts each flow on both ENIs and in both directions. Pick one direction, or one ENI per path, before summing.
  • SKIPDATA ignored. Totals during a capacity event are lower than reality. Count log_status = 'SKIPDATA' in any report that drives a decision.
  • Full scans. A query without a day filter over a year of 1-minute logs from a busy VPC can scan terabytes. Use projection, set Athena workgroup per-query scan limits, and consider Parquet.
  • Unbounded CloudWatch retention. Log groups default to never expire; flow logs are high-volume and this becomes the largest line on the logging bill.

Trade-offs

The trade-offs are coverage against cost and resolution against volume. Logging ALL traffic for every VPC at 1-minute aggregation gives the best forensic record and the biggest bill; logging REJECT only is cheap but cannot answer who talked to whom or why the egress bill doubled. A common compromise is ALL to S3 everywhere at the default 10-minute interval (Nitro instances get 1-minute resolution anyway), with S3 lifecycle rules moving older data to colder storage, plus a small REJECT-only CloudWatch subscription for alarms. Extra fields cost little in storage and are impossible to add retroactively, so err toward including them.

What to do next

  1. List every VPC and confirm each has a flow log; enforce it with an AWS Config rule or an organisation-level deployment so new VPCs are covered automatically.
  2. Define one custom format that keeps the 14 default fields first and adds vpc-id, subnet-id, instance-id, tcp-flags, pkt-srcaddr, pkt-dstaddr, flow-direction and traffic-path.
  3. Deliver to a central S3 bucket, create the projected Athena table above, and run the rejected-traffic query to prove the pipeline end to end.
  4. Set retention: an S3 lifecycle policy and, for any CloudWatch subscription, an explicit log group retention period.
  5. Save the four queries as Athena named queries and rehearse the security group versus ACL diagnosis before you need it.
  6. Enable Route 53 Resolver query logging or GuardDuty to cover the DNS traffic flow logs cannot see.

Related reading: Amazon VPC for the routing and security group model, Amazon Athena for query cost and tuning, CloudWatch in depth for Logs Insights and metric filters, and Amazon GuardDuty for managed threat detection built on flow logs and DNS logs.

Key takeaway: A flow record is per-ENI, per-window metadata: addresses, ports, protocol, counts and an accept or reject verdict, never payload. Choose the field list carefully because it cannot be changed later, keep the default fields first, and add pkt-srcaddr, flow-direction and traffic-path. Deliver to S3, query with a partition-projected Athena table that every query filters by day, remember the traffic that is never logged, and use the stateful versus stateless rule to separate security group rejections from network ACL ones.