AWS Certified Data Engineer - Associate (DEA-C01) study guide
4 min read · Updated July 20, 2026 · AI-assisted, editorially reviewed
How to use this guide
Give each domain study time in proportion to its official weighting. The percentages below are where your marks actually come from. Work through the domains in order, then let your per-domain practice scores show you what to go back to.
The fastest way to improve: read one domain’s focus areas, take a practice run, then read the explanation for every question you missed before moving on. The explanations are where the learning happens. Skipping them is the most common reason people stop improving.
1. Data Ingestion and Transformation
34% of the examWhat to focus on:
- Streaming and batch ingestion with Kinesis, MSK, DMS, Glue and Lambda
- Transformation with Glue ETL, EMR/Spark and SQL-based pipelines
- Orchestration, scheduling, retries and idempotency across pipeline stages
- Handling late-arriving, malformed and skewed data without silent loss
2. Data Store Management
26% of the examWhat to focus on:
- Choosing the store to fit the access pattern: S3, Redshift, DynamoDB, RDS
- Partitioning, file formats and compression, and what they cost at query time
- Schema design and evolution, including the Glue Data Catalog
- Lifecycle, tiering and retention across a data lake
3. Data Operations and Support
22% of the examWhat to focus on:
- Monitoring pipelines and diagnosing failures from CloudWatch and job logs
- Performance tuning: distribution and sort keys, partition pruning, concurrency
- Data quality checks, and validating that a pipeline produced what it claims
- Cost control across storage, scan volume and compute
4. Data Security and Governance
18% of the examWhat to focus on:
- IAM and Lake Formation permissions, including row and column level control
- Encryption with KMS, key rotation, and protecting data in transit and at rest
- PII handling: masking, tokenisation, and classification
- Auditing, lineage and meeting retention or residency requirements
Common mistakes
- Studying services rather than access patterns. The exam repeatedly asks which store or format fits a stated query and cost profile.
- Under-weighting Data Ingestion and Transformation, which alone is 34% of the exam.
- Ignoring file format and partitioning economics. Parquet versus JSON, and partition pruning, decide a surprising number of questions.
- Treating security and governance as an afterthought. Lake Formation, KMS and PII handling are 18% of the exam.
- Assuming 720 means 72%. It is a scaled score, and reasoning backwards from a percentage will mislead your readiness estimate.
Exam-day tips
- 130 minutes for 65 questions is about two minutes each, which is comfortable if you do not stall on the long pipeline scenarios.
- Read the constraint before the options. Latency, cost ceiling, retention period or compliance regime usually decides between two workable designs.
- Multiple-response items state how many to select; partial selections earn nothing.
- Prefer the managed or serverless option unless the scenario gives an explicit reason to self-manage.
- 15 of the 65 questions are unscored and indistinguishable, so do not let one baffling item shake your pacing.
How to know you’re ready
One good practice score can be luck. What counts is scoring at or above the real pass mark (720 on a 100 to 1000 scaled score) across several full-length sets in a row, with no single domain trailing far behind the others. Kwizza tracks your per-domain readiness automatically as you practice.
Free, full-length, and weighted to the official blueprint. Every answer is explained.
Start the Data Engineer practice exam →