Skip to content

Support bounded streaming body reads for sitemap acquisition #147

Description

@dacharyc

Motivation

PR #140 adds a cumulative 50 MiB sitemap-body budget checked between completed responses. This stops further sitemap fetches but does not bound one response: the final body can overshoot, and an unusually large response can still consume substantial memory before the budget is checked. Failed partial reads do not currently expose an acquired-byte count to the walker.

This is a distinct HTTP-layer follow-up, not a claim that the between-response budget is broken or an expansion requirement for PR #140.

Design Scope

Investigate exposing bounded body reads through the existing HTTP abstraction. The client already has internal bounded streaming for challenge inspection; reuse the appropriate machinery and preserve idle/total timeouts rather than adding a second independent reader.

Define whether the ceiling applies to decoded bytes, wire bytes, accepted text, or combinations, and distinguish it from cumulative scan budgets (#145). Specify unavoidable transport/chunk buffering and cancellation overshoot instead of promising zero bytes beyond a threshold.

Account for eager canonical-origin rewriting, shared text/body reads, denied-response inspection, clients without readable streams, and encoded/decoded size differences. Content-Length is insufficient for enforcement, especially for compressed or chunked responses.

Acceptance Criteria

  • A large sitemap response is stopped before its complete body is buffered or parsed; partially read XML is not silently treated as a complete sitemap.
  • Cancellation releases the response reader/resources and preserves existing timeout/error behavior.
  • Limit termination is distinguishable from a server failure or tarpit in request evidence and Expose structured discovery completeness and fallback evidence in reports #141 completeness metadata. Define useful byte accounting for partial reads.
  • Tests cover chunked bodies without Content-Length, gzip/brotli transfer decoding, a single oversized chunk, exact-limit bodies, canonical rewriting, repeated body/text reads, and cancellation.
  • The public/custom-client compatibility story is documented, including cases where a hard streaming guarantee cannot be provided.
  • Existing checks retain their normal body-reading behavior unless an explicit reviewed limit applies. This does not change page-size scoring thresholds.
  • Request-count and cumulative limits continue to apply; no additional probes or content-type discovery requests are introduced.

References

#120, PR #140, #141, #145, #104, #122. Relevant implementation: src/http.ts, HTTP request/response types in src/types.ts, and sitemap acquisition in src/helpers/get-page-urls.ts.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions