Skip to content

Expose structured discovery completeness and fallback evidence in reports #141

Description

@dacharyc

Motivation

Follow-up to #120 and PR #140, surfaced while drafting Documentation at Scale. Keep this separate from the bounded-walk fix.

The Microsoft Learn reproduction now completes, but discovery stops at its byte budget and falls back to a single base page: /en-us/docs matched 0 of 329,018 examined URLs. A completed scan, a partial discovery set, and an intentionally curated sample are different outcomes. Consumers should not need to parse warning strings to distinguish them.

Existing discoverySources and the single-page scoring safeguards from #23/#73 provide useful context, but do not establish completeness. Build on those rather than introducing another implicit score cap.

Design and Implementation Scope

Define structured metadata for discovery and coverage operations: source and effective scope, completion/stop reason, observed URL counts, relevant limits, and fallback provenance. Audit and reuse existing fields; counts of examined URLs are not the total size of a site.

Represent intentional selection/refinement separately from incomplete acquisition. Reaching the end of a root document is not proof that omitted nested indexes, skipped gzip sources, fetch failures, collection caps, or a large index were fully explored. Decide how unknown completeness is represented. Curated and none runs must not be labeled incomplete merely because they intentionally bypass sample discovery.

Propagate metadata through public types, cached samples, coverage results, reports, and relevant text/scorecard/JSON presentation. Define backward compatibility and avoid contradictory sources of truth.

Acceptance Criteria

  • Budget exhaustion, source errors, normal termination, URL-count limits, omitted traversal, intentional selection, and base-page fallback have documented semantics.
  • Discovery and coverage retain their own scopes and stop reasons; one completed operation cannot erase partial evidence from another.
  • Reports expose partial or unknown evidence without claiming exhaustive or representative site coverage.
  • CLI and programmatic consumers do not need to match warning text; existing human-readable warnings remain useful.
  • Tests cover cached and uncached paths, independent coverage walks, and retention of successful URLs after truncation.
  • No per-discovered-page probes or increased traversal depth. No implicit score or exit-code changes; those policies need separate decisions.

References

#120, PR #140, #25, #23, #73. Design history: working-notes/page-discovery-notes.md, section Documentation at scale: design considerations (currently in the local documentation draft). This issue is self-contained; it does not depend on that draft being published.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions