Yesterday a Go binary called mova showed up on r/PromptEngineering. It does something most of us never check: it audits what context your coding agent is about to send to an LLM, before the call happens. Most of us wire up an agent, watch it work, and never look at the payload. This tool puts that payload on the table.
What’s new
The builder, u/1982_miguel, calls it “context sovereignty before inference.” You decide what context is allowed to reach the model, and the tool leaves evidence of that decision. It’s open source and runs as a single Go binary, so there’s no server to stand up and no account to create. The pipeline looks like this:
Focus (AST) → PII masking → token budget → egress gate (dry-run) → LLM → evidence
Each stage has a clear job. AST focus keeps the code structure that matters for the task and drops the rest. PII masking pseudonymizes anything that looks like personal data. The token budget caps what’s left. The egress gate runs as a dry-run, so you can see what would be blocked before anything leaves your machine. The evidence step writes down what happened.
The repo ships with fictional data and a Linux amd64 build, so you can reproduce the numbers yourself. That matters here. A lot of tools in this space ask you to trust a benchmark chart. This one hands you the input and lets you rerun it.
The twist
The governance step doesn’t use an LLM. It needs no API key, and it doesn’t ask a model to guess what’s safe to send. It’s a deterministic step that sits in front of the model. That means the same input gives you the same output every time, which is exactly what you want from something acting as a checkpoint. You can’t audit a guess, but you can audit a rule.
It’s also not a gateway and not a RAG tool, which is an unusual lane. A gateway sits on the network and proxies calls. RAG retrieves context to add to a prompt. mova does neither. It works on the context you already have and decides what deserves to go out.
Here’s the example from the post:
- Before governance: 20,014 tokens
- After governance: 7,153 tokens (about 64% less)
- AST focus alone: 20,101 down to 5,325 tokens
- 171 of 1,694 PII-candidate tokens were pseudonymized
- A smaller fictional repo went from 1,965 to 779 tokens (about 60% less)
Notice that AST focus alone does most of the heavy lifting. That tells you something practical: much of what agents send is code that has nothing to do with the question being asked. Whole files go in when a few function signatures would have done the job.
One honest note on those numbers. The cost figures are theoretical input-token estimates, not real API spend. Treat them as a measurement of how much context got trimmed, not as an invoice. Your actual savings depend on your model, your caching setup, and how often your agent re-sends the same material.
Mini-workflow to try it
- 🔧 Grab the repo from github.com/m1guel1982/mova-context and use the Linux amd64 build. Run it on a throwaway machine or a clean directory first, since you’re testing a new binary.
- 📏 Run the reproducible example:
mova run --count 02-pii-compliance-governance. It prints the final token count (7,153). If your number differs, that’s worth reporting to the author. - 🔍 Compare before and after. Look at what AST focus kept versus what it dropped, then check the PII masking and the dry-run egress gate. This is where you see what would have been blocked before anything leaves your machine. Pay attention to the masked values and ask yourself whether you’d have caught them by eye.
- 🧾 Look at the evidence output. That’s your record of what the model was allowed to see. If you ever have to answer “what did the model receive?” for a client or a compliance review, this is the artifact you’d point to.
Pro tips
Smaller context is only a win if the model still has what it needs. A commenter on the thread (u/investigatormaker) made this point well: pair the token reduction with a cross-file task and check whether the trimmed context still contains every piece the task requires. Cutting 64% means little if the one file you needed got cut with it. A simple test is to pick a bug that spans two or three files, run it with and without governance, and compare the answers side by side.
Once you have a baseline, try your own repo instead of the fictional one. Start with a project that has no real customer data in it, so you can judge the trimming without worrying about what leaks.
Also read the limitations before you trust it with anything real. PII masking is heuristic, and the author hasn’t measured precision or recall yet. That means a masked value is a good sign, but an unmasked one isn’t proof of safety. mova only controls context that passes through it (CLI, chat, MCP, HTTP), so anything your IDE sends directly is outside its reach. macOS, Windows and arm64 builds are cross-compiled but not yet validated on those machines. The author lists all of this up front, which is a good sign.
Your turn
The author is asking for blunt feedback. Do you control what your agent sends today? Would you deliberately send less context for the same task? Do you want proof of what the model actually received? Run the example, then go tell them what you found, including “this solves a problem I don’t have.” They said that’s useful too. 🏴☠
Frequently Asked Questions
Q: Does reducing context really hurt your agent’s ability to solve tasks?
That’s the key trade-off. mova shows exactly which files and symbols survived the filtering, so you can verify your agent still has what it needs. The examples showed 60, 64% token reduction while preserving AST-focused context, just test it with your actual tasks to find the right balance.
Q: Can I see what got filtered out and why?
Absolutely. mova generates an evidence report showing what was included, excluded, and masked at each stage (AST focus, PII detection, token budget). That transparency means you can audit the decisions and debug cases where something important gets dropped.
Q: Is oversharing context with coding agents actually a real problem?
The tool’s creator kept running into it, massive context being sent when only small parts mattered. You might not notice if you’re under token limits, but it affects costs and can cloud your agent’s focus. mova helps you spot whether this is happening in your workflows.
Q: How does mova’s PII masking work?
It uses heuristic detection to find and mask potential PII (API keys, personal info, etc.) before context hits the LLM. It’s not bulletproof yet, the author hasn’t formally measured precision and recall. Always check the evidence report to make sure your sensitive data is actually protected.
How do you control what context your coding agent sends to an LLM? I built a local tool to measure and audit it — looking for blunt feedback
by u/1982_miguel in PromptEngineering