A Rust CLI Strips Up to 85% of Your Codebase Tokens Before You Paste Anything

Fresh drop on r/PromptEngineering: an open-source Rust tool called repOx. It aims at a problem most of us ignore. The wording of your prompt matters less than the junk you bury it in.

What’s new

repOx is a sub-15ms command line tool with an interactive terminal UI. It packs a repository into a prompt that’s ready to paste. The author’s pitch is that even with 200k to 1M context windows, raw dumps waste tens of thousands of tokens on structure. The model also gets lost in the middle and starts making things up. Bigger windows did not fix that. They just gave you more room to make the mess.

Think about what a typical “paste my repo” dump looks like. A few hundred lines of config, a lockfile the size of a novella, vendored folders, test fixtures, and then the three files you actually care about, buried somewhere in the middle. The model has to read all of it with equal attention. That’s the problem repOx goes after.

Version 0.2.0 handles the basics:

  • Wraps files in Claude-friendly XML (a repository structure block plus a file block per path), fenced Markdown, or synthetic JSON tool-call trajectories for agent harnesses
  • Filters binaries, SVGs, sourcemaps, minified bundles, and .env secrets
  • Shows a live offline token bar in the TUI before you copy to the clipboard

That last one is easy to underrate. Seeing the token count move as you tick files on and off teaches you fast which folders are the heavy ones. Most people are surprised the first time they see a single generated file eat half the budget.

The twist

Outline mode (repox --outline) strips function bodies. It keeps only types, structs, traits, imports, and function signatures. The author claims a 75 to 85% token cut while keeping the structural context. That’s perfect for “help me plan the architecture” prompts, where you never needed 10,000 lines of implementation anyway. Picture asking “where should I add a caching layer?” The model needs to see which modules exist, what they expose, and how they connect. It does not need the loop inside a string parser.

Lockfile summarization (repox --summary-locks) turns a 30,000-token Cargo.lock or package-lock.json into a compact “package @ version” list. The claim is a 95%+ reduction. Nobody’s prompt got smarter from reading checksums. You still keep the one thing the model might actually use, which is which versions you’re on. That helps when you ask about a breaking change or a dependency conflict.

Try it: mini-workflow

  1. 🛠️ Install with cargo install repox-cli, or use the one-liner script in the repo.
  2. 🔍 Run repox -i to open the TUI and watch the token bar as you pick files.
  3. 🧱 For planning questions, add --outline so the model sees the skeleton, not the guts.
  4. 📦 Add --summary-locks so lockfiles shrink to a version manifest.
  5. 📋 Copy the result into your chat and ask your architecture question.

If the answer feels thin, don’t panic. Go back to step 2, untick one or two outlined files, and include their full bodies. Then ask again. You’re tuning, not committing.

Pro tips

Send implementation bodies only for the files you actually want changed. Outline everything else. A good split is full code for two or three files, outline for the rest of the module, and nothing for unrelated folders.

One commenter said the XML structure with file paths is what keeps their Claude sessions coherent after 200k tokens. Paths as tags give the model something to anchor on. When you reply later and say “in the auth file,” the model can point back to a named block instead of guessing.

Another commenter made a fair challenge: don’t trust the savings number alone. Run the same small repo in full and in outline form, ask both the same architecture question, and compare the answers. Signatures preserved does not automatically mean reasoning preserved. Do that test before you rely on it. Pick a question where you already know the right answer, so you can grade both replies honestly.

Keep a note of the token count for each run. After a few sessions you’ll know your own sweet spot, and that number is more useful than any benchmark in a README.

What’s coming in v0.3.0

The author is designing an AI video side too, and this part is still a plan, not a shipped feature:

  • A character-lock plus delta prompt compiler. A short locked block holds subject, lens, and lighting. A tiny per-scene block holds only camera movement and action.
  • A storyboard contact-sheet packer. It packs keyframes into one timestamped 4×4 grid instead of 16 separate images, for roughly 93% fewer vision tokens.
  • A sharpest tail-frame anchor. It scans the last half second of a clip for the least blurry frame, to start the next shot.

The same idea runs through all three: say the stable stuff once, then only send what changes. It’s the video version of outline mode.

The author is also asking the community two things. Which format gives you the best reasoning accuracy for big packed prompts: XML, Markdown, JSON, or custom delimiters? And how do you keep character consistency across AI video shots without bloating the prompt? Worth dropping your answers in the thread!

It’s free and MIT licensed: github.com/WVDYC/repOx

⭐ Try it on your own repo, run the full vs outline comparison, and tell the crew what you find. A star on GitHub helps the author too.

Frequently Asked Questions

Q: Does the –outline feature preserve enough context for architectural decisions, or does stripping function bodies lose critical information?

Great question. The short answer: –outline is designed for architectural *planning*, not implementation review. It preserves function signatures, types, imports, and trait definitions , which is usually 80% of what you need for “how is this codebase structured?” But it *will* miss dependencies hidden only inside function bodies. The best way to validate this for your repo is to test the same architecture question against both the full and outline versions, then compare answers. If you’re asking “will this refactor break anything?” you probably need the full context.

Q: What happens to dependencies that only appear inside function bodies when using –outline?

They get stripped. If a dependency is only imported and used inside a function body (not at the module level), –outline won’t catch it. This is a known trade-off: you save 75, 85% of tokens but lose some internal call chains. If internal dependencies matter for your prompt (e.g., debugging a specific function’s behavior), use the full context or a filtered version of the file instead.

Q: Why does the XML structure keep context coherent at 200k+ tokens, when other formats fall apart?

The structured XML wrapper (with file paths and clear delimiters) helps the model maintain spatial awareness of where each chunk came from. It’s less about the XML itself and more about preventing the “lost in the middle” effect , explicit boundaries and metadata act like signposts that keep the model from treating a massive dump as undifferentiated noise.

Q: When should I use –outline vs. full repo dump vs. lockfile summarization?

–outline for architecture planning and refactoring decisions. Full dump for bug investigation (you need every detail). Lockfile summarization when your prompt is for dependency analysis, security audits, or upgrade planning , not implementation.

How are you structuring massive context prompts (codebases & multi-shot AI video) without burning 80% of tokens on noise? Built an open-source Rust tool (repOx) and looking for feedback
by u/Wvdy_CC in PromptEngineering

Scroll to Top