Cursor Is the Most Correct Tool. So Why Am I Not Using It?
Cursor has the sharpest read on where software is going, and I'm still not using it. The model is the product; everything above it is rented.
I think Cursor figured out where this whole job is going. I’m also not using it right now. Let me try to explain that.
The Most Correct Tool in the Room
It’s open season on Cursor right now. The pricing saga alone could fill a postmortem: the Pro plan everyone treated as effectively unlimited switched from request limits to compute credits, the effective cost of heavy agent workflows jumped something like twenty-fold, refunds went out, and people started leaving for Windsurf and Claude Code. Then Cursor 3 arrived - rebuilt from scratch under the codename Glass - and did the one thing nobody expected from the company that won by building the best editor on the market: it demoted the editor. The New Stack put it plainly when it said the IDE is now a fallback rather than the default.
Here’s the part that gets me funny looks: I think it’s the most correct tool in the category. The Medium headline that kicked off the backlash - “Cursor 3 Is Not an IDE Update. It’s a Bet That You’ll Manage Agents, Not Write Code” - was written as an accusation, but I read it as the right thesis. Typing code is becoming commoditized, and the work is shifting from producing keystrokes to overseeing the agents that produce them. Cursor pointed the whole company in that direction before its competitors had finished arguing about autocomplete. They saw it first, which is the hardest kind of right to be, and I’m rooting for them.
Which leaves a paradox on the table. The most accurate read of the future belongs to the company everyone’s piling on - and I’ve quietly stepped away too. If the vision is right, and I think it is, then my reason for not using it can’t be the vision. So what is it?
So Why Am I Not Switching?
Most people leaving Cursor are leaving for reasons a spreadsheet could fix: they got burned on a bill or hated seeing the editor shoved into a side panel. Those are real complaints, but clearer pricing and a layout toggle could solve them in a quarter.
Mine aren’t on that list. I didn’t leave over price or layout. I’m the sympathetic case: I agree with the bet, I want the overseer future to show up, and I’d happily pay whoever nails it. When the customer nodding along to your roadmap still reaches for something else, that tells you something a price cut won’t fix. In my case, it points to two problems, and neither is that the vision is wrong.
Reason One: The Model Has Its Own Gravity
The first reason is simple: Opus 4.8 is the best model I can get my hands on. Once you’ve spent real time with a model that feels clearly smarter than the rest, going back feels like trading down a level. Interface polish doesn’t close that gap, because the gap is in the part that does the thinking.
We were sold the opposite story for years: the model is the swappable commodity, and the tool is the moat. You could pick an editor, plug in whichever model you liked, and let the ergonomics win your loyalty. That held up while the models were close enough that the difference was rounding error. It comes apart the moment one pulls ahead, because then the model is the reason the output is any good.
I call it model gravity. The best model bends everything toward itself, and the pull lives in the small stuff you stop noticing you asked for. Opus 4.8 finds the edge cases on its own - the empty state, the race condition I’d have forgotten - without me writing “handle the edges” anywhere. It builds better UIs by default, the kind I’d otherwise have to art-direct line by line. It writes code that already matches our paradigms instead of the generic textbook version I’d have to rewrite. The roadmap doesn’t write the code; the model does, and right now the best one for my work isn’t Cursor’s.
I’ve said publicly that the model isn’t the thing that matters, and I still believe it - long term these models converge and this entire reason evaporates. Every one of those saves is really a prompting problem. I could write the context that makes a weaker (and more affordable) model find the same edges and match the same paradigms. I just don’t, because Opus is far enough ahead that I don’t have to, and the pricing is fine, so the effort never clears the bar. That’s the honest shape of model gravity today: paying for the best model beats doing the prompting work to live without it. There’s nothing permanent about that. If Anthropic’s pricing stops working for me, I’ll improve our prompts and switch to a Qwen 3.7 Plus, MiniMax 3, Composer 2.5, or whatever comes next.
Reason Two: A Promising Model That Doesn’t Think Like One Yet
Now the fair part, because the easy thing here is a cheap shot and I don’t want to take it. Cursor builds its own model, Composer, and Composer 2.5 has real potential - I’m bullish on it. It’s fast, it’s tuned for exactly the workflows Cursor cares about, and owning your model end to end beats renting one. Building it is the right long-term call.
But potential isn’t readiness. In my daily use, Composer 2.5 didn’t behave like an agent yet. The tell was how it handled ambiguity: it kept trying to write deterministic code - script the procedure, run the procedure - instead of sitting with uncertainty, trying things, backing out of dead ends, and noticing when the spec itself was wrong. Add a pile of small quirks that each cost thirty seconds and together break your flow, and you get something good at generating code but not yet good at being an agent. I had to ask it multiple times to act like one, and it kept falling back to shell scripts for tasks that could not be solved deterministically.
That sounds academic until you remember what Cursor is selling. The whole bet assumes the thing you’re managing is autonomous enough to be worth managing. A model that defaults to deterministic codegen is quietly fighting its own company’s strategy - the cockpit is built for an autonomous aircraft, but the engine still wants to be driven like a car. Fast, fluent codegen is mostly a solved problem; reliable agentic reasoning is the actual frontier, and the big labs are throwing everything at it. Cursor is chasing that frontier while also building an editor, a cloud platform, a pricing model, and a whole interface paradigm. Of those priorities, the model is the one that can least afford divided attention. This paragraph might be wrong in six weeks, but it isn’t today.
The Pattern: Model and Loop Beat Vision
Put the two together and they rhyme: the best model isn’t theirs, and the model they own doesn’t think like an agent yet. Strip away the specifics and it becomes one complaint. The value lives in the model underneath and the review loop on top, while the vision of managing a fleet sits in neither place. That vision is the one thing Cursor got dead right, and it turned out to matter least.
That should change what everyone’s competing over, and mostly it hasn’t. The whole field is racing to build dispatch: spin up a sandbox, clone the repository, run the agent, open the PR, and fan out across a swarm. But dispatch is a commodity. The most honest line I read all year, in a thread full of people trying to define “async agents,” suggested that “background job” was the more accurate framing. That’s all dispatch is, and a job queue isn’t a moat.
The line from that same thread that actually matters is that “the supervision protocol is the product, not the async dispatch.” That’s the whole game. The unsolved, defensible layer is the loop where a human reviews and corrects what the agent did - the last real bottleneck, because output scaled to thousands of lines an hour and human reading speed didn’t. Whoever makes that loop trustworthy and fast owns the workflow. That loop also wants a smart model underneath it, one that can summarize its own diff and flag where it’s unsure, which favors whoever already has the best model. Notice where the people leaving Cursor actually land: mostly in terminals such as Claude Code and Codex. Those tools are winning on the model and the surface alone, and they haven’t even seriously built the review loop yet.
So I Roll My Own - Until I Change My Mind in Six Weeks
So I’m in a terminal, building my own version of Cursor: the best model wired to a supervision layer designed around exactly how I work. It’s lean, it’s mine, and it’s good at the one job I need it to do. Cursor’s vision is right; I just don’t want to rent it from them yet. I’d rather wire the best model to a review surface I control than rent a polished workbench whose model isn’t agentic yet and whose pricing I’ve stopped trusting. Call it a stopgap if you want - a placeholder until someone ships the product that doesn’t exist yet, with the best model behind a painless loop. But it’s a real solution, and by my own argument it already has the two things that matter: the model underneath and the loop on top. It skips only the polish that turned out to matter least.
There’s a quiet condition under all of this, though: building my own is cheap because I’m the one running it. It works because there’s one set of hands on it, and the setup effort is small enough to clear the bar. Scale it to a team and the math inverts. Now it needs onboarding, shared infrastructure, and someone to maintain it, and those hours can cost more than the model gap I’d close by renting. That’s the real tipping point for someone like me: not the vision becoming more correct, but the day doing it myself stops being the easy win. The other tipping point is the one I already named: the models converge, the gap evaporates, and the deciding factor becomes whoever wraps the best loop around them. Until one of those two days arrives, the homegrown version wins on the merits.
If there’s a lesson for Cursor here, it’s that being right about the future doesn’t help much if you don’t own the model under it - and that you bring your users along instead of leaving them in the present. The people who loved Cursor signed up for the fastest editor on earth and got asked to become fleet managers nearly overnight. Maybe the overseer paradigm deserved its own product line: a new thing for the people ready for it, sitting next to the editor everyone already paid for. You can be early and still get ahead of your own users, and trust doesn’t come back with a changelog.
Here’s the honest part, and I think it proves the rest: this whole verdict expires in about six weeks. I’m not loyal to Claude Code; I’m loyal to whatever model is best this month, and the title keeps changing hands. The day a better one drops - maybe Composer 3 finally thinks like an agent and becomes my new go-to - I’ll throw out everything I just said and follow it. If a single release can flip my model, terminal, and workflow, then the model is the product and everything above it is rented. That includes both the vision and my harness. The only thing worth getting good at is switching, while keeping the part that’s actually yours - your judgment and your taste for what “good” looks like - light enough to carry to whatever wins next.
To me, Cursor has described the future more clearly than anyone. I can’t wait to find out who executes on it best.