mirror of
https://github.com/diegosouzapw/OmniRoute.git
synced 2026-08-12 02:02:13 +03:00
Scaffold the standalone Rust compression core (no OmniRoute deps): - crates/core-api: stable traits (TokenCounter, Compressor) + types (Message, Encoding, CompressionConfig/Result) — the contract for all adapters (N-API, sidecar, CLI) - crates/tokenizer: tiktoken-rs cl100k_base + o200k_base port - crates/tests: golden tests reading fixtures/ (byte-equality vs JS) - crates/bench: criterion harness (21.4ms vs JS 37.9ms on large input) - crates/ffi: empty N-API adapter placeholder (integration phase) - scripts/generate-fixtures.ts: JS reference output (source of truth) - scripts/verify-golden.ts: regen + run golden tests - fixtures/: 13 tokenizer samples incl. 405K-char stress case Golden status: 100% match JS vs Rust on all fixtures.
compression-core
Standalone Rust core for AI context optimization — tokenization, compression, hashing, translation primitives. Independent OSS library usable by OmniRoute, OpenCode, Cline, Roo, and any AI proxy. No OmniRoute imports anywhere.
Layout
compression-core/
├── Cargo.toml # workspace
├── crates/
│ ├── core-api/ # stable public API (traits + types) — no host deps
│ ├── tokenizer/ # tiktoken (cl100k_base, o200k_base) — PORTED
│ ├── tests/ # golden tests against fixtures/expected/
│ ├── bench/ # criterion benchmarks
│ └── ffi/ # N-API adapter (integration phase)
├── fixtures/
│ ├── tokenizer/ # JS-generated token counts (13 samples)
│ └── expected/ # manifests
└── scripts/
├── generate-fixtures.ts # JS reference output (source of truth)
└── verify-golden.ts # regen + cargo test
Porting order (per design)
- tiktoken (done — golden 100%)
- ionizer
- headroom
- caveman
- RTK (last — biggest, requires proven harness)
Golden pipeline
fixtures → JS implementation → expected.json → Rust → assert_eq!
node scripts/verify-golden.ts regenerates fixtures from the current JS code
and runs cargo test -p compression-tests. Until 100% match, JS stays in prod.
Measured baseline
| Impl | Input | Cost |
|---|---|---|
| JS js-tiktoken (cl100k) | 230K chars | 37.9 ms |
| Rust tiktoken-rs (cl100k) | ~440K chars | 21.4 ms |
Per-char Rust is ~3x faster; golden output is byte-identical on all fixtures.
Status
- workspace + stable API (
core-api) - tokenizer port + golden tests (100% match)
- bench harness (criterion)
- ionizer
- headroom
- caveman
- RTK
- N-API adapter