Motivation
Follow-up to #120 and PR #140, surfaced while drafting Documentation at Scale. Keep this separate from the bounded-walk fix.
The Microsoft Learn reproduction now completes, but discovery stops at its byte budget and falls back to a single base page: /en-us/docs matched 0 of 329,018 examined URLs. A completed scan, a partial discovery set, and an intentionally curated sample are different outcomes. Consumers should not need to parse warning strings to distinguish them.
Existing discoverySources and the single-page scoring safeguards from #23/#73 provide useful context, but do not establish completeness. Build on those rather than introducing another implicit score cap.
Design and Implementation Scope
Define structured metadata for discovery and coverage operations: source and effective scope, completion/stop reason, observed URL counts, relevant limits, and fallback provenance. Audit and reuse existing fields; counts of examined URLs are not the total size of a site.
Represent intentional selection/refinement separately from incomplete acquisition. Reaching the end of a root document is not proof that omitted nested indexes, skipped gzip sources, fetch failures, collection caps, or a large index were fully explored. Decide how unknown completeness is represented. Curated and none runs must not be labeled incomplete merely because they intentionally bypass sample discovery.
Propagate metadata through public types, cached samples, coverage results, reports, and relevant text/scorecard/JSON presentation. Define backward compatibility and avoid contradictory sources of truth.
Acceptance Criteria
- Budget exhaustion, source errors, normal termination, URL-count limits, omitted traversal, intentional selection, and base-page fallback have documented semantics.
- Discovery and coverage retain their own scopes and stop reasons; one completed operation cannot erase partial evidence from another.
- Reports expose partial or unknown evidence without claiming exhaustive or representative site coverage.
- CLI and programmatic consumers do not need to match warning text; existing human-readable warnings remain useful.
- Tests cover cached and uncached paths, independent coverage walks, and retention of successful URLs after truncation.
- No per-discovered-page probes or increased traversal depth. No implicit score or exit-code changes; those policies need separate decisions.
References
#120, PR #140, #25, #23, #73. Design history: working-notes/page-discovery-notes.md, section Documentation at scale: design considerations (currently in the local documentation draft). This issue is self-contained; it does not depend on that draft being published.
Motivation
Follow-up to #120 and PR #140, surfaced while drafting Documentation at Scale. Keep this separate from the bounded-walk fix.
The Microsoft Learn reproduction now completes, but discovery stops at its byte budget and falls back to a single base page:
/en-us/docsmatched 0 of 329,018 examined URLs. A completed scan, a partial discovery set, and an intentionally curated sample are different outcomes. Consumers should not need to parse warning strings to distinguish them.Existing
discoverySourcesand the single-page scoring safeguards from #23/#73 provide useful context, but do not establish completeness. Build on those rather than introducing another implicit score cap.Design and Implementation Scope
Define structured metadata for discovery and coverage operations: source and effective scope, completion/stop reason, observed URL counts, relevant limits, and fallback provenance. Audit and reuse existing fields; counts of examined URLs are not the total size of a site.
Represent intentional selection/refinement separately from incomplete acquisition. Reaching the end of a root document is not proof that omitted nested indexes, skipped gzip sources, fetch failures, collection caps, or a large index were fully explored. Decide how unknown completeness is represented. Curated and
noneruns must not be labeled incomplete merely because they intentionally bypass sample discovery.Propagate metadata through public types, cached samples, coverage results, reports, and relevant text/scorecard/JSON presentation. Define backward compatibility and avoid contradictory sources of truth.
Acceptance Criteria
References
#120, PR #140, #25, #23, #73. Design history:
working-notes/page-discovery-notes.md, sectionDocumentation at scale: design considerations(currently in the local documentation draft). This issue is self-contained; it does not depend on that draft being published.