Capstone: AWS Inventory Reporter
Combine pagination, SDK retries, timeouts, logging, command-line options, CSV reports, and offline tests in one project.
Prerequisites
Complete the production scripting bridge, EC2 reporting, and retry labs. The offline project needs Python 3 and no third-party packages. Optional live mode needs boto3, configured credentials, and ec2:DescribeInstances permission in your chosen region.
Project brief
An operations team needs a repeatable EC2 inventory report. Build a command-line tool that:
- Reads either local response fixtures or all EC2 response pages in one region.
- Produces a deterministic CSV with region, ID, state, type, and name.
- Handles missing Name tags and empty inventories.
- Uses connection/read timeouts and at most four SDK attempts per request.
- Logs a count and region, and returns a nonzero exit code on failure.
- Refuses to overwrite an existing report and removes a partial report if writing fails.
- Can be tested without credentials or network calls.
Sample input and downloads
Save these files in the same directory. The IDs in the fixture are fictional.
- sample-pages.json — two pages with one instance each.
- test_inventory.py — behavior tests for your implementation.
Write your own inventory.py before opening the reference implementation below.
Expected output
python inventory.py --fixture sample-pages.json --region us-east-1 --output inventory.csv
Stdout: Wrote 2 instances to inventory.csv. The file contains:
region,id,state,type,name
us-east-1,i-example1,running,t3.micro,web-01
us-east-1,i-example2,stopped,t3.micro,unnamed
The region flag labels offline fixture records; fixtures themselves contain no region metadata. A second run using the same output path must fail without changing the first report. Use a new path for the next successful run.
Hint: break the project into four parts
Separaterows_from_pages(pages, region), collect_live(client, region), write_report(rows, output), and main(argv=None). Inject a fake client into the collection function. Import boto3 only in live mode so offline use needs no SDK installation.Show the reference solution
Download inventory.py. It uses the SDK paginator, bounded standard retries, explicit timeouts, a pure normalization function, and exclusive output-file creation. It makes no AWS calls unless you explicitly pass --live.
The report currently collects and sorts all rows in memory. For a very large fleet, choose a streaming or external-sort design. Pagination alone does not bound report memory usage.
Run the tests
python -m unittest -v test_inventory.py
Tests cover multiple pages, missing tags, stable ordering, empty output, malformed input, output preservation, partial-write cleanup, and propagated pagination errors. A mocked paginator exercises the collection contract; it does not simulate the SDK retry engine.
Optional live run
Install boto3 in your virtual environment and use your configured AWS profile or role:
python -m pip install boto3
python inventory.py --live --region us-east-1 --output live-inventory.csv
Add --profile training if you use a named profile. This mode reads EC2 inventory and does not create or modify AWS resources. Permission or credential failures must fail the run; they must not become a successful empty report.
The SDK handles supported transient errors within the configured attempt budget. See SDK retries and the EC2 paginator. A whole-job deadline, distributed rate limits, and multi-account credentials are outside this first version.
Completion checklist
- Generate the exact sample report and pass the offline tests.
- Explain why only reading the first response page loses data.
- Demonstrate a failed run leaves no new report.
- Explain timeouts, total attempts, and failure exit codes.
- Give a five-minute walkthrough of your design and one limitation.
Further challenges
Add multiple regions, a missing-Owner-tag report, and a scheduled CI job that publishes the CSV as an artifact. Test one failed region and decide whether partial results should count as success. If users open reports in spreadsheet software, add and test a policy for untrusted tag values that resemble formulas.
Continue to production scenarios or the interview revision route.
Add More Questions to This Guide
Know a question that should be here? Share it and help the community!
Open Google Form