I have been experimenting with local large language models for coding for a while. OpenCode was an important part of that journey because it gave me a practical way to connect different providers, run local models, and learn where those models were genuinely useful. It worked well, but I kept wanting a harness that felt more robust during long sessions and made it easier to give different kinds of work to different models.
Credit for the discovery goes to a coworker who mentioned OMP, or Oh My Pi, during a conversation at the office. I had no idea this harness existed, and that casual recommendation sent me down a very productive rabbit hole. After using OMP on real projects, this is the best agentic coding setup I have configured so far. It feels fast, steady, and reliable enough that I can focus on the work instead of constantly thinking about the harness. This is still an early impression rather than a final verdict, but OMP has already helped me turn a collection of AI tools into a coherent development workflow.

Local efficiency and cloud reasoning meet at one deliberate human review point.
The harness matters as much as the model
A coding model never works alone. The harness decides which files it can inspect, how edits are applied, what tools are available, how context is preserved, and whether the model can use capabilities such as language servers, debuggers, browsers, and subagents. That surrounding system can significantly change the quality of the result, even when the underlying model stays the same.
OMP is a coding-agent harness with a substantial native Rust core, built as a coding-focused fork of Pi. It supports hosted and local models, but the provider list is not what won me over. The interesting part is the engineering around the model: hash-anchored edits that reject stale references, Language Server Protocol integration for code intelligence, Debug Adapter Protocol integration, persistent Python and JavaScript environments, structured subagents, browser tools, and dedicated review and planning workflows.
The result feels closer to giving a model a carefully designed engineering environment than simply placing a chat interface inside a terminal. File reads are focused, searches are quick, and edits have safeguards that reduce the cycle of failed replacements and repeated attempts. Rust is part of the performance story, but the more important point is that OMP treats the tool layer as a product rather than an afterthought.
One model does not need to do everything
It is easy to treat model selection as a winner-takes-all decision. Choose the strongest model available, send every request to it, and accept the subscription or API cost. That is convenient, but it is not always efficient. Planning a large refactor or investigating an architectural problem may justify the strongest reasoning available. Exploring files, updating isolated code, or implementing a clearly bounded task often does not.
OMP addresses this with model roles. The default role handles ordinary work, smol favors quick and inexpensive work, slow is intended for deeper reasoning, and plan powers planning sessions. Other roles cover delegated tasks, vision, commits, and advisors. Instead of choosing a model again for every prompt, I can decide what each model is good at and let the harness preserve that division of labor.
My approach is simple: Codex handles difficult reasoning and planning, while Qwen3.6 handles suitable coding work on my own machine. OMP sits between them, and I remain responsible for defining the boundaries and reviewing the result. The model is a worker selected for the job, not the architecture of the entire workflow.
My local sweet spot
My desktop has an Intel Core i5-13600K, 32 GB of system memory, and an AMD Radeon RX 7900 XTX with 24 GB of VRAM. It is not a datacenter, but it is a very capable local inference machine. With an appropriate quantization and context size, I can run Qwen3.6 locally and give it meaningful coding tasks without paying for every token.
There are real constraints. The model needs to fit within the available memory budget, longer contexts can reduce performance, and a local open model will not match a frontier model on every difficult reasoning problem. Running locally also means owning more of the setup, including the serving software, drivers, model selection, quantization, and performance tuning. In return, the model is available whenever I need it, it does not consume subscription quota, and the marginal cost of another task is effectively the electricity required to run it.
When a problem needs stronger reasoning, I can move it to Codex through the subscription I already have. This is not about proving that local models are better than cloud models, or the reverse. It is about not paying for the most expensive intelligence on every problem. On my hardware, that balance is the practical sweet spot.
Where OMP fits beside OpenCode
OpenCode is still a strong project. It has a broad provider ecosystem, supports local models, and offers terminal, desktop, and editor experiences. It helped me establish the local workflow that eventually led me to OMP. The difference is not that OpenCode is bad and OMP is good. They emphasize different things.

OpenCode emphasizes broad access, provider-native agents emphasize a direct ecosystem, and OMP emphasizes a deep tool harness with explicit model routing.
My choice
OMP
A deep coding harness built around explicit model roles, strong code intelligence, and deliberate routing between local and cloud models.
Best fit: hybrid local and cloud engineering
Broad access
OpenCode
An approachable multi-provider experience with excellent model choice across terminal, desktop, and editor workflows.
Best fit: exploring models and providers
Direct ecosystem
Provider-native agents
A focused path optimized around one provider, its models, and the conventions of its own product ecosystem.
Best fit: convenience within one ecosystem
This is a summary of my priorities, not a benchmark or a universal ranking. OpenCode made model choice easy for me. OMP goes further by helping me turn that choice into an operating strategy. I am not merely switching models inside the same interface. I am assigning each model a job and giving it a toolset designed for software work.
How I divide the work
The configuration is the easy part. The useful skill is deciding where each task belongs. I send bounded and easy-to-verify work to local Qwen: repository exploration, module summaries, repetitive updates, isolated implementation steps, and independent subagent tasks. These jobs benefit from speed and availability, and they usually do not need my limited cloud allowance.
Codex handles work where weak reasoning would be more expensive than additional usage. That includes planning changes across several systems, evaluating architectural tradeoffs, diagnosing subtle failures, reviewing sensitive code, and challenging an implementation before it is merged. I am not saving Codex only for emergencies. I am spending its stronger reasoning where it creates the most leverage.
Some decisions remain mine regardless of which model is active. I decide what the product should do, whether a requirement is correct, which operational risks are acceptable, and whether generated code belongs in the system. OMP can coordinate models and tools, but it does not remove engineering responsibility.

The efficient path is not local or cloud. It is deliberate routing followed by human review.
What I am still evaluating
My experience has been excellent so far, but I still want more time with long sessions, larger repositories, context compaction, and local subagents. I also want to measure how much subscription capacity this hybrid setup saves in practice and learn which of OMP’s advanced capabilities become part of my daily workflow rather than features I only admire on paper.
For now, I am deliberately keeping the process simple. Strong planning goes to Codex. Suitable implementation goes to local Qwen. Every result comes back through my review. That simplicity is one reason the setup feels stable.
The best setup I have used so far
Local models gave me control and a very low marginal cost. Hosted frontier models gave me stronger reasoning. Until now, combining the two often felt like maintaining separate workflows. OMP is the first harness I have used that makes the combination feel intentional.
On my 13600K, 32 GB of RAM, and Radeon 7900 XTX, Qwen3.6 gives me a capable local worker. My Codex subscription provides the deeper reasoning I want for planning and difficult decisions. OMP connects them without forcing me to abandon either side. The result is the best balance I have found between capability, speed, reliability, and cost.
I am not ready to call this the final answer. Agentic development tools are changing quickly, and the difficult edge cases only appear with time. I will keep testing OMP and plan to return with a longer-term review once the novelty has worn off. For now, it is staying in my terminal.