Skip to content

Sum DynamoDB ConsumedCapacity across paginated pages (CLI-4199) - #10690

Open
GarrettBeatty wants to merge 6 commits into
v2from
cli-4199-wildcard-paginators
Open

GarrettBeatty wants to merge 6 commits into
v2from
cli-4199-wildcard-paginators

Conversation

@GarrettBeatty

@GarrettBeatty GarrettBeatty commented Sep 24, 2026 •

Copy link
Copy Markdown

Issue

aws dynamodb scan/query with auto-pagination (the default) reports only the final page's ConsumedCapacity instead of the sum across all pages, undercounting consumed capacity. (CLI-4199 / D103095602)

Concrete example

A scan that spans two pages against a table with a global secondary index, e.g. page 1 consumes 448.5 CU and page 2 consumes 63.5 CU (512.0 total):

aws dynamodb scan --table-name Orders --index-name byCustomer \
    --return-consumed-capacity INDEXES

Before — the aggregated result carries only the last page's numbers:

{
    "Items": [ /* all items from both pages */ ],
    "Count": 1000,
    "ScannedCount": 1000,
    "ConsumedCapacity": {
        "TableName": "Orders",
        "CapacityUnits": 63.5,                          // last page only
        "Table": { "CapacityUnits": 0.0 },
        "GlobalSecondaryIndexes": {
            "byCustomer": { "CapacityUnits": 63.5 }     // last page only
        }
    }
}

After — every numeric leaf is summed across pages, each index keeps its own counter, and TableName comes from the first page:

{
    "Items": [ /* all items from both pages */ ],
    "Count": 1000,
    "ScannedCount": 1000,
    "ConsumedCapacity": {
        "TableName": "Orders",
        "CapacityUnits": 512.0,                         // 448.5 + 63.5
        "Table": { "CapacityUnits": 0.0 },
        "GlobalSecondaryIndexes": {
            "byCustomer": { "CapacityUnits": 512.0 }    // summed per index
        }
    }
}

With --no-paginate (single page) output is unchanged. When --return-consumed-capacity is not requested, ConsumedCapacity is absent from both the responses and the result (no phantom {"TableName": null}).

Approach

The paginator config keys understand wildcard (*) path segments:

  • result_key — a path may contain * segments. Each concrete leaf matched on every page is summed (numeric leaves only), and each distinct matched map-key keeps its own running counter. This is what reaches the runtime-keyed per-index maps (GlobalSecondaryIndexes, LocalSecondaryIndexes, VectorIndexes) that plain dotted paths can't express.
  • non_aggregate_keys — a wildcard/dotted path takes its value from the first page (e.g. ConsumedCapacity.TableName).

A * is only honored when the service model confirms the *'s parent member is a map shape (and, for result keys, the leaf resolves to a numeric shape). Otherwise the wildcard path is dropped (it is never compiled as a plain jmespath key, where * would become a list projection that gets silently concatenated across pages) — so a service can't accidentally opt into wildcard aggregation just by shipping a * in its model. A @validates_models allowlist in the paginator-config test enforces which (service, operation) pairs may use wildcard result keys.

DynamoDB config (paginators-1.sdk-extras.json overlay)

For Scan and Query, result_key adds the top-level ConsumedCapacity.{CapacityUnits,ReadCapacityUnits,WriteCapacityUnits}, ConsumedCapacity.Table.*, and the per-index maps ConsumedCapacity.{GlobalSecondaryIndexes,LocalSecondaryIndexes,VectorIndexes}.*.*; non_aggregate_keys adds ConsumedCapacity.TableName.

Engine changes (awscli/botocore/paginate.py)

  • _iter_wildcard_leaves(data, segments) — yields (concrete_path, value) for each leaf present in a page matching the segment pattern (* matches every dict key at that level).
  • _walk_to_leaf_container / _is_summable_number — helpers to accumulate into the result and to guard against booleans / non-numerics.
  • build_full_result sums wildcard numeric leaves per page; _record_non_aggregate_key_values records wildcard leaves from the first page.
  • A nested non_aggregate leaf that is None on the first page is skipped only when its parent object is itself absent from the response (recording it would fabricate a phantom parent like {"TableName": null} implying the member is present). If the parent is present, a genuinely-null leaf is still surfaced as null — so services with nullable nested non_aggregate keys (e.g. kinesis DescribeStream's StreamDescription.KeyId, athena GetQueryResults) keep their exact prior output.
  • Paginator._is_map_wildcard_path performs the model-shape gate; result/non-aggregate keys are split into plain (compiled jmespath, unchanged) vs. validated-wildcard lists. An all-wildcard result_key (no plain primary key) raises a clear PaginationError instead of a later IndexError.

Tests

  • Unit (tests/unit/botocore/test_paginate.py): wildcard leaf iteration, per-key summation across pages, the map-shape gate + fallback (built against a real ServiceModel), nested null-leaf preserved when parent present, phantom parent skipped when parent absent, gate-failing wildcard dropped, and all-wildcard result_key raises.
  • Functional (tests/functional/dynamodb/test_pagination.py): scan/query total sum, per-GSI wildcard sum with distinct counters, VectorIndexes sum, and no ConsumedCapacity (not even a phantom {"TableName": null}) when capacity isn't requested.
  • Paginator-config validation: first result_key may not be a wildcard; wildcard numeric map leaves require the allowlist.
  • Verified end-to-end against live DynamoDB (on-demand table + GSI, multi-page scan).

Implements the ticket design (doc L9lObjjWpDcq): wildcard ('*') paths in the
existing result_key (numeric, summed) and non_aggregate_keys (first page)
rather than a new directive. A '*' matches every key of a map member and each
matched shape gets its own counter.

- paginate.py: PageIterator sums wildcard numeric result-key leaves per matched
  map key across pages and records wildcard non-aggregate leaves from the first
  page; Paginator gates '*' on the model and falls back otherwise.
- dynamodb overlay: ConsumedCapacity numeric leaves in result_key;
  ConsumedCapacity.TableName in non_aggregate_keys.
- linter + tests + changelog.
Adversarial review round 1: moving non_aggregate_keys to ConsumedCapacity.TableName
makes the paginator materialize ConsumedCapacity as {TableName: null} when
capacity isn't requested. The ddb strip only removed a None ConsumedCapacity, so
the phantom object leaked (broke 24 functional ddb tests). Drop ConsumedCapacity
when it is None or an all-None dict. (Also modernizes pre-existing super() calls
flagged by ruff.)
- Standard 'aws dynamodb scan/query' also leaked a phantom
  ConsumedCapacity:{TableName:null} when capacity wasn't requested (round-1
  only fixed the 'aws ddb' path). Stop materializing non_aggregate leaves whose
  value is None in the paginator, so an absent member (e.g.
  ConsumedCapacity.TableName) is simply omitted on every path instead of
  surfacing null. This also makes plain non_aggregate handling consistent with
  wildcard non_aggregate (which only records present matches).
- Add ConsumedCapacity.VectorIndexes.*.{VectorSearchRequestBytes,
  VectorWriteRequestBytes} wildcard result keys so vector-index capacity is
  summed rather than dropped when paginating.
- Tests: assert ConsumedCapacity fully absent when not requested; add
  VectorIndexes aggregation test.
…nested paths

The round-2 None-skip was too broad: it dropped null-valued top-level
non_aggregate members for all 26 services that use non_aggregate_keys (e.g.
kinesis describe-stream KeyId). Scope the skip to NESTED paths only: a nested
path (ConsumedCapacity.TableName) would fabricate a phantom parent object
{TableName: null}, so skip it; a top-level scalar member still surfaces as null
(its historical behavior) since that doesn't fabricate a parent. Fixes the
DynamoDB not-requested phantom on all paths while preserving cross-service
output for the common top-level case.
…wildcard config, drop dead code

- non_aggregate None-skip now only skips a nested null leaf when its PARENT
  is absent from the response (would fabricate a phantom parent). A null leaf
  on a present parent is preserved, fixing a regression for kinesis
  DescribeStream (StreamDescription.KeyId) and athena GetQueryResults.
- _get_result_keys / _get_non_aggregate_keys no longer fall through to
  jmespath.compile for a wildcard that fails the map gate (where '*' becomes a
  list projection silently concatenated across pages); such paths are dropped
  with a debug log.
- _get_result_keys raises a clear PaginationError when result_key yields no
  non-wildcard primary key, instead of a later IndexError in __iter__.
- Removed unreachable elif branch in the wildcard sum and the now-redundant
  ConsumedCapacity strip in ddb subcommands (phantom is fixed at the source).
- Tests for null-leaf-kept, phantom-parent-skipped, wildcard-dropped, and
  all-wildcard-raises.
@GarrettBeatty
GarrettBeatty marked this pull request as ready for review September 28, 2026 18:49
@GarrettBeatty
GarrettBeatty requested a review from a team as a code owner September 28, 2026 18:49
_handle_first_request truncates secondary result keys on a starting-token
resume and previously wrote None back for any absent key. With the new nested
ConsumedCapacity.* result keys this fabricated a phantom
ConsumedCapacity={CapacityUnits: None, ...} in the page, which then made
_record_non_aggregate_key_values see a present parent and add TableName: None
(surfacing ConsumedCapacity even when --return-consumed-capacity was not set).

Skip writing an absent (None) secondary result key instead of fabricating it.
Adds a functional regression test for scan --starting-token without capacity.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant