Skip to content
#

byte-level

Here are 18 public repositories matching this topic...

A tiny byte-level multi-head content classifier (~1.5M params, ~200KB ONNX, <6ms). Classifies code, text, markup, config, images, binary, secrets, 62 code languages, 30 text languages, 90 MIME types from raw bytes — no tokenizer needed.

  • Updated Sep 26, 2026
  • Python
purebyte

Tiny byte-level AI on CPU: zero-dependency C++20 inference engine and neural decision models (2.8ms latency, 9.4MB RAM). Replaces tokenizers with a 256-byte vocabulary. Flagship specialists for pre-commit secret scanning (0.797 F1 vs 0.337 GitLeaks) and PII redaction. Apache 2.0.

  • Updated Sep 28, 2026
  • C++
purebyte-train

The official training stack for PureByte: train byte-level neural decision specialists from a task definition to a verified GGUF model in ~25 minutes on a single GPU. Includes synthetic hard-negative generation, curriculum training, and automated evaluation against CredData and PIIMB.

  • Updated Sep 28, 2026
  • Python

Add this topic to your repo

To associate your repository with the byte-level topic, visit your repo's landing page and select "manage topics."

Learn more