Sum DynamoDB ConsumedCapacity across paginated pages (CLI-4199) - #10690
Open
GarrettBeatty wants to merge 6 commits into
Open
GarrettBeatty wants to merge 6 commits into
GarrettBeatty wants to merge 6 commits into
Conversation
Implements the ticket design (doc L9lObjjWpDcq): wildcard ('*') paths in the
existing result_key (numeric, summed) and non_aggregate_keys (first page)
rather than a new directive. A '*' matches every key of a map member and each
matched shape gets its own counter.
- paginate.py: PageIterator sums wildcard numeric result-key leaves per matched
map key across pages and records wildcard non-aggregate leaves from the first
page; Paginator gates '*' on the model and falls back otherwise.
- dynamodb overlay: ConsumedCapacity numeric leaves in result_key;
ConsumedCapacity.TableName in non_aggregate_keys.
- linter + tests + changelog.
Adversarial review round 1: moving non_aggregate_keys to ConsumedCapacity.TableName
makes the paginator materialize ConsumedCapacity as {TableName: null} when
capacity isn't requested. The ddb strip only removed a None ConsumedCapacity, so
the phantom object leaked (broke 24 functional ddb tests). Drop ConsumedCapacity
when it is None or an all-None dict. (Also modernizes pre-existing super() calls
flagged by ruff.)
- Standard 'aws dynamodb scan/query' also leaked a phantom
ConsumedCapacity:{TableName:null} when capacity wasn't requested (round-1
only fixed the 'aws ddb' path). Stop materializing non_aggregate leaves whose
value is None in the paginator, so an absent member (e.g.
ConsumedCapacity.TableName) is simply omitted on every path instead of
surfacing null. This also makes plain non_aggregate handling consistent with
wildcard non_aggregate (which only records present matches).
- Add ConsumedCapacity.VectorIndexes.*.{VectorSearchRequestBytes,
VectorWriteRequestBytes} wildcard result keys so vector-index capacity is
summed rather than dropped when paginating.
- Tests: assert ConsumedCapacity fully absent when not requested; add
VectorIndexes aggregation test.
…nested paths
The round-2 None-skip was too broad: it dropped null-valued top-level
non_aggregate members for all 26 services that use non_aggregate_keys (e.g.
kinesis describe-stream KeyId). Scope the skip to NESTED paths only: a nested
path (ConsumedCapacity.TableName) would fabricate a phantom parent object
{TableName: null}, so skip it; a top-level scalar member still surfaces as null
(its historical behavior) since that doesn't fabricate a parent. Fixes the
DynamoDB not-requested phantom on all paths while preserving cross-service
output for the common top-level case.
…wildcard config, drop dead code - non_aggregate None-skip now only skips a nested null leaf when its PARENT is absent from the response (would fabricate a phantom parent). A null leaf on a present parent is preserved, fixing a regression for kinesis DescribeStream (StreamDescription.KeyId) and athena GetQueryResults. - _get_result_keys / _get_non_aggregate_keys no longer fall through to jmespath.compile for a wildcard that fails the map gate (where '*' becomes a list projection silently concatenated across pages); such paths are dropped with a debug log. - _get_result_keys raises a clear PaginationError when result_key yields no non-wildcard primary key, instead of a later IndexError in __iter__. - Removed unreachable elif branch in the wildcard sum and the now-redundant ConsumedCapacity strip in ddb subcommands (phantom is fixed at the source). - Tests for null-leaf-kept, phantom-parent-skipped, wildcard-dropped, and all-wildcard-raises.
_handle_first_request truncates secondary result keys on a starting-token
resume and previously wrote None back for any absent key. With the new nested
ConsumedCapacity.* result keys this fabricated a phantom
ConsumedCapacity={CapacityUnits: None, ...} in the page, which then made
_record_non_aggregate_key_values see a present parent and add TableName: None
(surfacing ConsumedCapacity even when --return-consumed-capacity was not set).
Skip writing an absent (None) secondary result key instead of fabricating it.
Adds a functional regression test for scan --starting-token without capacity.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Issue
aws dynamodb scan/querywith auto-pagination (the default) reports only the final page'sConsumedCapacityinstead of the sum across all pages, undercounting consumed capacity. (CLI-4199 / D103095602)Concrete example
A
scanthat spans two pages against a table with a global secondary index, e.g. page 1 consumes 448.5 CU and page 2 consumes 63.5 CU (512.0 total):Before — the aggregated result carries only the last page's numbers:
{ "Items": [ /* all items from both pages */ ], "Count": 1000, "ScannedCount": 1000, "ConsumedCapacity": { "TableName": "Orders", "CapacityUnits": 63.5, // last page only "Table": { "CapacityUnits": 0.0 }, "GlobalSecondaryIndexes": { "byCustomer": { "CapacityUnits": 63.5 } // last page only } } }After — every numeric leaf is summed across pages, each index keeps its own counter, and
TableNamecomes from the first page:{ "Items": [ /* all items from both pages */ ], "Count": 1000, "ScannedCount": 1000, "ConsumedCapacity": { "TableName": "Orders", "CapacityUnits": 512.0, // 448.5 + 63.5 "Table": { "CapacityUnits": 0.0 }, "GlobalSecondaryIndexes": { "byCustomer": { "CapacityUnits": 512.0 } // summed per index } } }With
--no-paginate(single page) output is unchanged. When--return-consumed-capacityis not requested,ConsumedCapacityis absent from both the responses and the result (no phantom{"TableName": null}).Approach
The paginator config keys understand wildcard (
*) path segments:result_key— a path may contain*segments. Each concrete leaf matched on every page is summed (numeric leaves only), and each distinct matched map-key keeps its own running counter. This is what reaches the runtime-keyed per-index maps (GlobalSecondaryIndexes,LocalSecondaryIndexes,VectorIndexes) that plain dotted paths can't express.non_aggregate_keys— a wildcard/dotted path takes its value from the first page (e.g.ConsumedCapacity.TableName).A
*is only honored when the service model confirms the*'s parent member is amapshape (and, for result keys, the leaf resolves to a numeric shape). Otherwise the wildcard path is dropped (it is never compiled as a plain jmespath key, where*would become a list projection that gets silently concatenated across pages) — so a service can't accidentally opt into wildcard aggregation just by shipping a*in its model. A@validates_modelsallowlist in the paginator-config test enforces which(service, operation)pairs may use wildcard result keys.DynamoDB config (
paginators-1.sdk-extras.jsonoverlay)For
ScanandQuery,result_keyadds the top-levelConsumedCapacity.{CapacityUnits,ReadCapacityUnits,WriteCapacityUnits},ConsumedCapacity.Table.*, and the per-index mapsConsumedCapacity.{GlobalSecondaryIndexes,LocalSecondaryIndexes,VectorIndexes}.*.*;non_aggregate_keysaddsConsumedCapacity.TableName.Engine changes (
awscli/botocore/paginate.py)_iter_wildcard_leaves(data, segments)— yields(concrete_path, value)for each leaf present in a page matching the segment pattern (*matches every dict key at that level)._walk_to_leaf_container/_is_summable_number— helpers to accumulate into the result and to guard against booleans / non-numerics.build_full_resultsums wildcard numeric leaves per page;_record_non_aggregate_key_valuesrecords wildcard leaves from the first page.Noneon the first page is skipped only when its parent object is itself absent from the response (recording it would fabricate a phantom parent like{"TableName": null}implying the member is present). If the parent is present, a genuinely-null leaf is still surfaced asnull— so services with nullable nested non_aggregate keys (e.g. kinesisDescribeStream'sStreamDescription.KeyId, athenaGetQueryResults) keep their exact prior output.Paginator._is_map_wildcard_pathperforms the model-shape gate; result/non-aggregate keys are split into plain (compiled jmespath, unchanged) vs. validated-wildcard lists. An all-wildcardresult_key(no plain primary key) raises a clearPaginationErrorinstead of a laterIndexError.Tests
tests/unit/botocore/test_paginate.py): wildcard leaf iteration, per-key summation across pages, the map-shape gate + fallback (built against a realServiceModel), nested null-leaf preserved when parent present, phantom parent skipped when parent absent, gate-failing wildcard dropped, and all-wildcardresult_keyraises.tests/functional/dynamodb/test_pagination.py): scan/query total sum, per-GSI wildcard sum with distinct counters, VectorIndexes sum, and noConsumedCapacity(not even a phantom{"TableName": null}) when capacity isn't requested.