Skip to content
@QuesmaOrg

Quesma

Making AI agents production-ready through independent evaluation and training.

Pinned Loading

  1. CompileBench CompileBench Public

    Benchmark of LLMs on real open-source projects against dependency hell, legacy toolchains, and complex build systems.

    Astro 60 8

  2. otel-bench otel-bench Public

    OpenTelemetry Benchmark - can AI trace your failed login?

    Shell 23 4

Repositories

Showing 10 of 24 repositories
  • quesma-shipper Public

    Collects AI coding-agent sessions from developer machines, scrubs secrets, encrypts with age, and uploads to your organisation's storage.

    QuesmaOrg/quesma-shipper's past year of commit activity
    Go 19 Apache-2.0 1 3 9 Updated Sep 29, 2026
  • homebrew-tap Public
    QuesmaOrg/homebrew-tap's past year of commit activity
    Python 0 0 0 0 Updated Sep 28, 2026
  • awesome-ai-tokenomics Public

    A curated list on AI token economics: what tokens cost, where they get wasted, and how to cut the bill. Tools, benchmarks, papers, and copy-paste configs for the token economy of LLMs and coding agents.

    QuesmaOrg/awesome-ai-tokenomics's past year of commit activity
    Python 190 CC0-1.0 34 0 1 Updated Sep 25, 2026
  • shipper-protocol Public

    The wire contract between Quesma Shipper and its control plane.

    QuesmaOrg/shipper-protocol's past year of commit activity
    Go 6 Apache-2.0 0 0 1 Updated Sep 16, 2026
  • terminal-bench-science Public Forked from harbor-framework/terminal-bench-science

    Terminal Bench for Science

    QuesmaOrg/terminal-bench-science's past year of commit activity
    Shell 1 Apache-2.0 381 0 0 Updated Jul 25, 2026
  • otel-bench Public

    OpenTelemetry Benchmark - can AI trace your failed login?

    QuesmaOrg/otel-bench's past year of commit activity
    Shell 23 Apache-2.0 4 0 4 Updated Jul 14, 2026
  • BinaryAudit Public

    An open-source benchmark for evaluating AI agents' ability to find backdoors hidden in compiled binaries.

    QuesmaOrg/BinaryAudit's past year of commit activity
    Shell 101 7 3 1 Updated Jul 14, 2026
  • QuesmaOrg/trival-prompt-bench's past year of commit activity
    HTML 1 0 0 1 Updated Jul 14, 2026
  • CompileBench Public

    Benchmark of LLMs on real open-source projects against dependency hell, legacy toolchains, and complex build systems.

    QuesmaOrg/CompileBench's past year of commit activity
    Astro 60 MIT 8 2 3 Updated Jul 14, 2026
  • gradient-engineer Public

    Gradient Engineer: 60‑Second Linux Analysis (Nix + LLM)

    QuesmaOrg/gradient-engineer's past year of commit activity
    Go 18 MIT 3 0 1 Updated Jul 3, 2026