# Cursor (Anysphere)

Feedback Bench, coding agents, built 2026-10-01, window 2026-08-31 to 2026-09-27. Web page: https://feedbackbench.com/#/agent/cursor

| Measure | Value |
|---|---|
| Rank | 4 of 17 (rank range 4–4) |
| Feedback Score | 59.7 (95% interval 59.0–60.5) |
| Popularity | 0.708 (share of voice 8.74%) |
| Customer love | 0.504 (95% interval 0.491–0.516) |
| Top quadrant | yes, at the edge |
| Authors | 8609 |
| Posts counted | 18878 |
| Posts that judge the agent | 9985 |
| Criteria better / worse than peers | 5 / 10 of 63 |

## The brief

Written by Claude Opus 5.5 from 134 labelled posts and the numbers on this page. Interpretation, not measurement: every quote is verbatim and links to its post.

**Strong harness and IDE, undercut by support silence and opaque quotas.**

TL;DR:

- Doing the work is a strength; users credit task capability, plan mode and the IDE.
- Support is the sharpest pain: tickets loop through a bot and billing cases wait unanswered.
- Auto routing and the Grok push make users unsure which model, and quota, they are spending.

### What hurts

- **Support loops through a bot** ([Support, refunds and issue handling](https://feedbackbench.com/criteria/account.support.md), [Wrong charges, failed payments and plan provisioning](https://feedbackbench.com/criteria/account.billing_errors.md)). Account and billing problems stall because escalation routes users back to an AI assistant, and duplicate-ticket handling closes their follow-ups.
  Users describe emailing support and getting the same bot that answers on Reddit. One sign-in case went in circles. Emails from the blocked account were closed as duplicates, and the help portal needed the login that was broken.
  
  Billing cases fare no better. Posts report days of silence on charges users do not recognise. One user says a disputed bill topped $2,000.
  Evidence:
  - Complaint, Cursor, r/cursor, 2026-09-03: “i emailed them and i just got the same bot that answers here. this is crazy. i had so much work” [source](https://www.reddit.com/r/cursor/comments/1vzv0id/cursor_banned_my_account_after_15_years_with_zero/p7oe03i/)
  - Complaint, Cursor, @cursor_ai, 2026-09-08: “@cursor_ai your system sucks. i've been trying to use @bot for days and still can't get access. i'm going in circles with sam, your ai assistant. i just want to sign in with my x sso account but keep getting "access blocked." i've emailed, used the forums, and asked @grok, but nothing escalates to a human. emails from the blocked account get closed as duplicates instead of being added to the original thread. even replies on the original email are closed as duplicates. i can't use the help portal because i can't sign in with that email address. emails from another address trigger a security issue and i'm told by sam the ai to reply to the existing email, which then gets closed as a duplicate. i’m a supergrok subscriber and should be able to use this service. help! @elonmusk” [source](https://twitter.com/1892892127136534528/status/2097215559720894798)
  - Complaint, Cursor, r/cursor, 2026-09-21: “i contacted cursor support regarding a serious billing and account security issue and opened a support ticket. i provided the requested information and followed up with additional details, but i haven't received a response or meaningful update on the investigation. the issue involves more than **$2,000 in charges**, including usage that i don't recognize, so i'm concerned about waiting without knowing the status of the case. i'm not trying to resolve the entire billing dispute through reddit. i mainly want to know: has anyone recently had a similar experience waiting for cursor support? how long did it take before you received a response on a billing or account-security case? if anyone from the cursor team sees this, i'd really appreciate someone checking my existing support case. i can provide the ticket number privately. thanks.” [source](https://www.reddit.com/r/cursor/comments/1wm37db/is_cursor_support_currently_taking_a_long_time_to/)
  - Complaint, Cursor, @cursor_ai, 2026-09-23: “@xai @cursor_ai @grok actually canceled gemini/claude for @poteto's cursor credits, bought $300 grok sub too was supposed to get the cursor/grok deal xai support no response for 37 days since august 17 cursor support no response for 5 calendar days / 3 business days, billed $200” [source](https://twitter.com/1450247589631336448/status/2102799122923131301)

- **Stalls and slowdowns with no word** ([Latency, throughput and fast mode](https://feedbackbench.com/criteria/rel.response_speed.md), [Outages, server errors and capacity or rate errors](https://feedbackbench.com/criteria/rel.service_errors.md), [Client crashes, freezes and failed tool execution](https://feedbackbench.com/criteria/rel.client_failures.md)). Requests hang and need retries, users suspect shared model capacity, and the official explanation they want does not arrive while it happens.
  Posts describe Composer stalling even on fast mode and Grok running noticeably slower, with requests retried repeatedly. One user reports a fraction of normal speed and wonders whether capacity moved to a new model. The agents window draws its own complaint for slow startup.
  
  The contrast stings. Other users praise Composer for staying up when every other model was down.
  Evidence:
  - Complaint, Cursor, r/cursor, 2026-09-08: “yes, same here since mid last week. composer stalls constantly even on fast, and grok has been noticeably slower too. i've had to retry requests multiple times because it just hangs. feels like they're rate limiting or the servers are overloaded, but no official word from cursor yet” [source](https://www.reddit.com/r/cursor/comments/1wam12r/composer_and_grok_slow_down_and_stalling/p8jwwtp/)
  - Complaint, Cursor, r/cursor, 2026-09-13: “yes, i am probably running at 4% of normal speed. i was curious if i did something that caused it to slow. i also had a huge bug introduced... it may be because of a substantial change but i am curious if xai is putting all resources toward the new model at the cost of model performance.” [source](https://www.reddit.com/r/cursor/comments/1wexjrt/anyone_else_experiencing_slow_down_right_now/p9ho9wo/)
  - Complaint, Cursor, r/cursor, 2026-09-05: “100% agree, it makes no sense for most people. go check the comments under the antigravity 2.0 launch video, where they tried pulling the same thing. on top of this, i don’t know what they used for the cursor agents window but the startup time is offensively slow on my computer” [source](https://www.reddit.com/r/cursor/comments/1w7wdar/the_agents_window_is_cursor_telling_you_to_stop/p7y7lb2/)
  - Complaint, Cursor, @cursor_ai, 2026-09-13: “grok 4.6 is practically unusable today. @cursor_ai @grok please <strict_link>” [source](https://twitter.com/1814045931349946368/status/2099213864525005203)

- **Auto mode can quietly downgrade you** ([Automatic model routing and fallback](https://feedbackbench.com/criteria/models.routing_auto.md), [Quality got worse or better over time](https://feedbackbench.com/criteria/models.quality_drift.md)). Auto routing falls back to weaker models without telling users, so they only notice when output quality drops, and an upstream Grok outage takes Auto down with it.
  One user says Auto dropped them onto a weaker model and they only caught it after a review went soft. Their fix is to pin models and pick fallbacks by hand. Others note Auto is effectively Composer or Grok, so when Grok is down, Auto is down. Users also report plans executed by a cheaper model that missed what was asked. No silent model switching is a standing request.
  Evidence:
  - Complaint, Cursor, r/cursor, 2026-09-14: “same afternoon squeeze here. i never flip to auto when grok is busy. auto has silently dropped me onto a weaker model before and i only noticed after the review got soft. if the pinned model is rejected or overloaded i pick the nearest equivalent tier myself or wait it out. silent fallback is how you lose the model you thought you were running.” [source](https://www.reddit.com/r/cursor/comments/1wfxcf0/grok_is_extremely_busy/p9prqd1/)
  - Complaint, Cursor, r/cursor, 2026-09-03: “auto is obviously composer or grok. grok down takes the whole thing down. oops” [source](https://www.reddit.com/r/cursor/comments/1w67nlp/is_auto_so_locked_on_to_grok_that_it_wont_even/p7l04sh/)
  - Complaint, Cursor, r/cursor, 2026-09-21: “so mine defaulted to this, i had grok 4.7 xhigh plan out a few tasks, they were executed with composer 2.5. it missed a lot in the plan, and didnt do what i asked. this was across a handful of plans.” [source](https://www.reddit.com/r/cursor/comments/1wmm80e/grok_47_xhigh_plan_composer_25_build_fast_mode_off/pb8f7my/)
  - Complaint, Cursor, @cursor_ai, 2026-09-02: “@cursor_ai so composer is now in a dead spot, always more expensive than luna, and always worse than grok? didn't realize that, need to update my defaults...” [source](https://twitter.com/1024209146/status/2095257403524600216)

- **Quota drains faster than the meter shows** ([Single prompt, model or effort level consumes disproportionate quota](https://feedbackbench.com/criteria/limits.burn_rate.md), [Usage meter visibility and accuracy](https://feedbackbench.com/criteria/limits.usage_meter.md), [Pay-as-you-go overage, fallback billing and spend caps](https://feedbackbench.com/criteria/billing.overage_charges.md), [Quota reset timing and bonus or banked resets](https://feedbackbench.com/criteria/limits.reset_schedule.md)). Single prompts and heavy models can consume most of a monthly allowance, and the usage meter is too buried and unclear for users to see it coming.
  Users report one prompt eating the remainder of an Ultra month. The meter does not help. People ask to see dollars spent inside the chat, complain they must open settings to check what is left, and one found usage hitting 100% from models they say they never chose. Overage complaints follow the same line, with users saying extra cost arrives while the meter stays hidden.
  Evidence:
  - Complaint, Cursor, @cursor_ai, 2026-09-16: “@benvargas @cursor_ai i burnt through remaining 73% of my monthly cursor ultra with one fable prompt” [source](https://twitter.com/1727559851184652289/status/2100172557282427266)
  - Complaint, Cursor, @cursor_ai, 2026-09-14: “@cursor_ai i have noticed at times the main usage of all models i had around 97% done and i never used anything on frontier model but it reached 100% i checked the exports but i just see random models being used what is this? the export data doesnt not show it right as now it has way too much mix with grok bot shown as well in cursor dashboard anyone else had this issue?” [source](https://twitter.com/2242568539/status/2099627906666516760)
  - Complaint, Cursor, @cursor_ai, 2026-09-20: “hey @cursor_ai it would be helpful to see the current spending in dollars directly within the chat.” [source](https://twitter.com/308597676/status/2101601519217033245)
  - Complaint, Cursor, @cursor_ai, 2026-09-10: “@mgallmur @cursor_ai @claudeai yesss, exactly. they add an extra cost while hiding the meter” [source](https://twitter.com/1251095647345754115/status/2098143814909370560)

- **Thin outside the desktop app** ([Mobile, remote-control and voice access](https://feedbackbench.com/criteria/surfaces.remote_mobile.md), [Connecting own API keys, local models and custom endpoints](https://feedbackbench.com/criteria/setup.provider_byok_local.md)). Projects and agent history live mainly in the desktop app, so mobile, web and bring-your-own-key users hit missing features or broken saves.
  Users who liked Projects found they could not see them or their agent history on mobile or web. A native Android app is a top request. Bring-your-own-key users fare worse than peers. One reports an API key that clears on every request. Another would pay a small fee for BYOK alone. iOS Projects shipped during the window, and users noticed.
  Evidence:
  - Complaint, Cursor, @cursor_ai, 2026-09-11: “@cursor_ai already start using and love this feature! but it seems only supported in desktop app, but not on mobile nor web, right? i could not see my projects nor agent running history on other devices” [source](https://twitter.com/1773662247023468545/status/2098247684914389355)
  - Complaint, Cursor, @cursor_ai, 2026-09-04: “@cursor_ai where is the android app? perhaps your team needs to use claude to build it.” [source](https://twitter.com/1572038422373568512/status/2095924424499154991)
  - Complaint, Cursor, @cursor_ai, 2026-09-06: “hello @cursor_ai disappointing to see basic bugs like this at this stage. i'm trying to use my own api key, but it won't save. it fails and clears out on every single request. is a fix in the works? <strict_link>” [source](https://twitter.com/1414879245948669953/status/2096619373389554164)
  - Complaint, Cursor, r/cursor, 2026-09-24: “coursor harness is the best in my opinion but the models are just too expensive now. i would subscribe for 5 bucks a month if they just gave me byok option.” [source](https://www.reddit.com/r/cursor/comments/1wpd3mm/i_hate_to_admit_it_but_grok_sucks/pbvao4e/)

### What works

- **Gets real tasks done, fast** ([Can do the user's kind of task](https://feedbackbench.com/criteria/work.capability.md)). Users who compare harnesses credit Cursor with speed and with solving problems other tools could not, which puts task capability better than peers.
  One user who tried several rival agents calls Cursor the fastest in any situation. A long-time user says it solved firmware issues no other model could crack. Web developers say Cursor's own models are much quicker for simple tasks and save frontier models for harder ones. The praise centres on the harness and iteration speed, not one model.
  Evidence:
  - Praise, Cursor, r/cursor, 2026-09-19: “after trying codex, devin, antigravity, claude, cursor is by far the fastest in any situation in terms of quality, im really enjoying gpt sol, just.. works, and fixes, anything” [source](https://www.reddit.com/r/cursor/comments/1wkr3d8/im_starting_to_think_they_pushed_some_update/pau0k22/)
  - Praise, Cursor, r/cursor, 2026-09-26: “to the mods and the community: i’d like to request that we keep this subreddit focused solely on cursor. i’m not here to be a blind advocate for the tool, but i’ve been using it since the very early days. back when i was building firmware, no other llms could figure out my issues, but cursor actually solved them. lately, there’s been a lot of noise about deepseek, opencode, and alternative ways to develop. even if those other methods might arguably be "better," that’s not what this sub is for. i want to come here to learn more about what's under the hood of cursor and how to improve our workflows inside it. can we enforce a rule to keep the discussions strictly on-topic? i am not sure if this sub reddit was made by the cursor itself. but if it is i request you to hire me for a pay or hire someone to strictly mod it idrc but do it.” [source](https://www.reddit.com/r/cursor/comments/1wqr30h/can_we_keep_this_subreddit_to_only_discuss_about/)
  - Praise, Cursor, r/cursor, 2026-09-27: “for my web development usage i find cursor's models like composer to be much faster for simple tasks. i also like the inbuilt browser and ide interface, despite how it seems like they are trying to make it an afterthought in the app. opus 5.5 is next level but i've only found myself reaching for that for tougher tasks.” [source](https://www.reddit.com/r/cursor/comments/1wrx24u/1_year_of_cursor_switched_to_claude_best_decision/pcgnlh4/)
  - Praise, Cursor, @cursor_ai, 2026-09-16: “@sheherenow_ this is why i loved @cursor_ai composer it was so fast iteration flow states were like the default” [source](https://twitter.com/1554700708225507328/status/2100306031817748842)

- **The IDE is the reason to stay** ([IDE and editor integration](https://feedbackbench.com/criteria/setup.ide_integration.md)). Users who want to read and edit code themselves pick Cursor for its IDE, which keeps integration better than peers even among critics of its models.
  Posts frame the IDE as the moat. People want to see the code and the architecture, and nudge UI by hand rather than prompt for it. Some pair the harness with models from several providers inside one editor.
  
  The agent view draws complaints for lacking full file search and language tooling. Some users run plain VS Code alongside.
  Evidence:
  - Praise, Cursor, r/cursor, 2026-09-18: “in your case it doesn’t make sense. i use cursor because i mostly use grok models and sometimes i use claude and gpt. this fits well in the two polls cursor provides. besides that, i like the ide features and harness. if you already pay for claude/codex it doesn’t make sense” [source](https://www.reddit.com/r/cursor/comments/1wjcyla/got_cursor_start_just_to_see_if_cursor_can_fit_my/pahy8k7/)
  - Praise, Cursor, @cursor_ai, 2026-09-17: “the only reason i use @cursor_ai is because it has an ide yes, you can vibe code stuff, but i want to see the code yes, it may write better than me, but i want to understand the architecture and where it all stored old fashion? heretic? no, just worked enough in high-load projects where ai still make silly mistakes that can crash the product like quering pg instead of ch also, it's way faster to me to play with ui via css than ask llm many times e.g "can you move the text to the right in footer" vs "text-right" i'm typing in both cases but the second one is way shorter” [source](https://twitter.com/1525392577612206080/status/2100529587100774757)
  - Complaint, Cursor, @cursor_ai, 2026-09-08: “one thing that bugged me in @cursor_ai is that in agentic view (glass as it is called internally) i do not have full support of file search, quick code jump and all other .net goodies incl .razor files edit. so i created my own version, implemented integration with .net tooling and voila, have it all now and am happy bunny. hey @cursor_ai if you want to know how to, let me know, will give you the tools” [source](https://twitter.com/36333025/status/2097307459480084979)
  - Complaint, Cursor, r/cursor, 2026-09-03: “cursor's "vs code ide with ai agent integration" is the shittiest shit i've ever seen. i use cursor in an agent view next to true vs code (not shitty one, without shitty redesign, without unnecessary features and with all vs code extensions) highly recommend this approach: cursor is for generating code, vs code is for working with the code. do not mix this” [source](https://www.reddit.com/r/cursor/comments/1w5nl3p/stop_forcing_me_to_the_agent_view_by_default/p7jx0m8/)

- **Plan first, then let it run** ([Plan-before-edit mode](https://feedbackbench.com/criteria/work.plan_mode.md)). Plan mode cuts wasted runs by surfacing the approach before edits, and users who adopt it say they wish they had started sooner.
  Users describe a sequence: discuss a change with /ask, review it with /plan, then let the agent execute, because misunderstandings waste time and quota. A two-year user calls planning mode the biggest difference. The rough edges show up in Projects. After long sessions, users say plans stop opening and build can loop.
  Evidence:
  - Praise, Cursor, @cursor_ai, 2026-09-03: “with @cursor_ai i am getting the best results from: - using /ask first to discuss the change - using /plan to see how it will do it then i let it rip. misunderstandings waste time and resources.” [source](https://twitter.com/271616165/status/2095382909188214826)
  - Praise, Cursor, r/cursor, 2026-09-10: “opus 4.6 is my daily driver in planning mode and to execute the plan, then when the task is complete if i have any follows ups i use got 5.2, not codex version. i find 5.2 works for longer and doesn't stop to ask you questions as much saving on time and context by skipping useless uodate messages that wait for a response. honestly if you're not using planning mode you should try it. been using cursor for about 2yrs and i wish i'd started using it sooner. makes a big difference.” [source](https://www.reddit.com/r/cursor/comments/1wc95mb/i_am_tired_of_handholding_composer_25_what_are/p8w6c0a/)
  - Praise, Cursor, @cursor_ai, 2026-09-12: “@fatih @cursor_ai created my first plan with it and started to implement it. thanks for the write-up 🫡 i will definitely keep using plan-add for all ideas that pop into my mind before they disappear for good” [source](https://twitter.com/1289301690315874305/status/2098765350070329390)
  - Complaint, Cursor, @cursor_ai, 2026-09-12: “@cursor_ai when in project mode, after an extended period of time you can no longer open plans from the “view plans” button when asking it to make a new one. this causes the plan to just be called “plan” and makes everything break. trying to hit build on this will cause loops” [source](https://twitter.com/2205898980/status/2098841482690167070)

- **One subscription, many models** ([How much use a plan's price buys](https://feedbackbench.com/criteria/limits.plan_value.md), [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md)). Users who sequence models carefully say a single plan stretches far, because cheap first-party models handle routine work while frontier models handle hard subtasks.
  Posts report shipping full client websites on the entry plan. Power users instruct a cloud coordinator to use Composer for most coding and Grok only for complex judgement, and say the subscription goes much further. Access to many models from one sub is a stated draw. The catch is that the praise depends on users managing routing themselves.
  Evidence:
  - Praise, Cursor, r/cursor, 2026-09-07: “i can totally agree, i can get up and running a website (frontend and backend) for a client whilst on the 20$ plan.” [source](https://www.reddit.com/r/cursor/comments/1w9udb4/is_cursor_worth_it/p8dvvyh/)
  - Praise, Cursor, r/cursor, 2026-09-23: “no. but i give my cursor projects agent (the one who orchestrated the cloud agents) specific instructions about what models to use. for example if i'm out of credits to use some of the frontier models, i tell it to use composer as the primary coding agent, only use grok 4.6, or now 4.7, high (ne er xhigh) when it's a particularly complex subtask or if judgment is required during the development. and i tell it to never use fast. the cloud coordinator ends up using composer about 80% of the time and i'm not switching back and forth between models. you should try projects and give it some specific instructions with a properly formed task and i think you'll be amazed at how far your subscription will go.” [source](https://www.reddit.com/r/cursor/comments/1wnpgmf/canceled_cursor_today_after_using_it_for_many/pbjpx6q/)
  - Praise, Cursor, r/cursor, 2026-09-03: “if you use 90% claude models, you should be on a claude sub for sure. for me, cursor’s usp is the ide (for how much longer, i dunno). first party models are genuinely good workhorses and - right now at least - access to loads of models from the one sub (until openai shuts the door). i make heavy use of bugbot but equally the claude code review skills are excellent (though i don’t think you can plug them into github actions without an enterprise plan) if the vscode extensions weren’t such a second rate offering, i’d probably be on a claude or openai sub only” [source](https://www.reddit.com/r/cursor/comments/1w6b34p/cursor_claude_code_codex/p7mjlyg/)
  - Praise, Cursor, r/cursor, 2026-09-18: “came from claude, used codex as well, but the cursor models provide me sufficient tokens for all i need. mainly use grok 4.5, composer and 4.6. the work that they provide is good. don’t recognize your comment.” [source](https://www.reddit.com/r/cursor/comments/1wjcyla/got_cursor_start_just_to_see_if_cursor_can_fit_my/paj03rt/)

### Under the surface

- **The Grok push reshapes the catalog** ([Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md), [Quality got worse or better over time](https://feedbackbench.com/criteria/models.quality_drift.md), [Automatic model routing and fallback](https://feedbackbench.com/criteria/models.routing_auto.md)). Users report losing OpenAI models and being steered to Grok, so complaints about catalog, quality and routing converge on one question of model choice.
  Posts say OpenAI models no longer run, that some users were moved to Grok involuntarily, and that cheaper options like Luna will leave with them. Reactions split. Some rate Grok's code quality, speed and low usage highly. Others want Composer 3 as a cheap, fast alternative, and it ties for the most-requested change. Expect sentiment to track the next first-party model release.
  Evidence:
  - Complaint, Cursor, @cursor_ai, 2026-09-24: “@stefanjblos @cursor_ai i can’t even run openai models on cursor now :/” [source](https://twitter.com/2382240745/status/2103083873676706107)
  - Praise, Cursor, @cursor_ai, 2026-09-24: “@e_viki_ @cursor_ai to be clear, i didn't move to grok intentionally, cursor moved me to grok involuntarily after the spacex purchase. with that said, code quality is actually pretty good. plan following and just following directions in general seems to be its weekest point.” [source](https://twitter.com/14311446/status/2102969871437033755)
  - Complaint, Cursor, r/cursor, 2026-09-22: “with the way cursor forces grok harder and harder on everyone - it will soon become not an option not to use it. i want composer 3. something cheap and fast, maybe not as smart as grok, but good enough for the majority of tasks.” [source](https://www.reddit.com/r/cursor/comments/1wmhp5w/grok_47_is_out_try_it/pbboj5u/)
  - Praise, Cursor, r/cursor, 2026-09-22: “i actually have a bit of an opposite view, yes you are right that grok is a bit meh but for a huge portion of tasks it's sufficient and it's really fast and costs little usage even in fast mode, for something like setting something up or something very boring, or making a quick/temporary code change in a codebase you don't care about, you can end up using up your entire claude/codex usage, meanwhile grok can do it extremely fast and use up practically no usage. i keep cursor grok for tasks like this meanwhile keeping the harder tasks for codex/claude. it's also quite lenient in terms of cyber security tasks and isn't as strict compared to codex/claude. you don't get much api usage from the $20 plan but it's a bonus for me personally considering grok makes up for the plan for me. personally i found composer to consistently be very bad, grok is sort of composer on steroids for me.” [source](https://www.reddit.com/r/cursor/comments/1wnpgmf/canceled_cursor_today_after_using_it_for_many/pbgwxy9/)

- **Parallel agents multiply the bill** ([Subagents, parallel agents and orchestrators](https://feedbackbench.com/criteria/work.multi_agent_orchestration.md), [Single prompt, model or effort level consumes disproportionate quota](https://feedbackbench.com/criteria/limits.burn_rate.md), [Cloud and remote sandbox execution](https://feedbackbench.com/criteria/surfaces.cloud_sessions.md)). Multi-agent and cloud fan-out is a draw, but orphaned or oversized subagents drain quota where users cannot watch it.
  Users like /multitask and shared memory for running a small team of agents at once. The same feature produces the worst burn stories. One user lost an entire monthly allowance for other models within minutes of reset. Others ask for an inspectable queue, and for killing a parent agent to also stop its children.
  Evidence:
  - Complaint, Cursor, @cursor_ai, 2026-09-18: “not only that... cloud agents multitasking burned my entire monthly "other model" credits in under *9* minutes at reset... 9 minutes. spun up max fable 5.1 agents x 24 (i had one agent set at medium, and had requested auto agents for the work with the fable orchestrator). *burned* never again.” [source](https://twitter.com/779746/status/2100770463685497087)
  - Praise, Cursor, @cursor_ai, 2026-09-15: “i love building with @cursor_ai - opus 5 for planning - gpt-5.6 sol for reasoning and debugging - grok 4.6 for the implementation. through /multitask, one agent can work on the feature, another can inspect the codebase, another can handle tests or debug. all working parallelly. you basically turn one coding session into a small team of agents working at the same time.” [source](https://twitter.com/1455270436090957826/status/2099837406547693994)
  - Complaint, Cursor, @cursor_ai, 2026-09-11: “@fazedordecodigo @peyton_nowlin @cursor_ai orphaned subagents are the expensive failure mode, not the fan-out itself, quota drains from work nobody is watching anymore. kill needs to walk the process tree on the same signal, not just close the parent's stdout.” [source](https://twitter.com/1447584564239486980/status/2098460120170778765)
  - Complaint, Cursor, @cursor_ai, 2026-09-10: “@cursor_ai @bot always-on only helps if the queue is inspectable. if i can't see what got assigned, what's blocked, and what done means per subagent, you just turned the babysitting into a longer chat.” [source](https://twitter.com/2819971425/status/2098185903697203630)

### Fine print

- Nearly all posts come from X replies tagging the vendor and r/cursor, which skew toward vocal, engaged users.
- Some criticism tracks ownership politics rather than product behaviour, as users themselves note, so scores may mix both.
- Trustpilot and G2 contribute only a handful of posts, and many narrow criteria rest on small samples.

## Top requests

What users ask to add or change, most asked first. 1806 author-weeks ask for something. Requests do not change the Feedback Score. Rule: A separate pass by Claude Sonnet 5 reads every counted post and extracts what the author asks the agent or its vendor to add or change, with the criteria it maps to and a short normalised wording; it does not touch the labels or the Feedback Score. Claude Opus 5.5 groups the wordings within each criterion (the first criterion the request maps to) into themes; code counts them. A theme counts distinct author-weeks that ask for it, per agent; across agents, one author-week per agent. Themes asked in fewer than 2 author-weeks, and requests that share no theme, are not shown. Examples: up to 3 posts per theme from different authors, without slurs, preferring posts of 60 to 450 characters, most recent first.

| Rank | Request | Criterion | Author-weeks | Posts |
|---|---|---|---|---|
| 1 | Native Android app | [Mobile, remote-control and voice access](https://feedbackbench.com/criteria/surfaces.remote_mobile.md) | 36 | 50 |
| 2 | Release Composer 3 model | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 36 | 37 |
| 3 | One-off usage limit reset now | [Quota reset timing and bonus or banked resets](https://feedbackbench.com/criteria/limits.reset_schedule.md) | 35 | 37 |
| 4 | Unified subscription across linked products | [Using an existing subscription across tools](https://feedbackbench.com/criteria/billing.subscription_portability.md) | 27 | 28 |
| 5 | Add GPT-6 Astra model | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 26 | 27 |
| 6 | Add Grok 4.7 model | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 24 | 24 |
| 7 | Higher overall usage limits | [How much use a plan's price buys](https://feedbackbench.com/criteria/limits.plan_value.md) | 19 | 19 |
| 8 | No silent model switching or downgrades | [Automatic model routing and fallback](https://feedbackbench.com/criteria/models.routing_auto.md) | 18 | 21 |
| 9 | Additional or recurring bonus usage resets | [Quota reset timing and bonus or banked resets](https://feedbackbench.com/criteria/limits.reset_schedule.md) | 18 | 19 |
| 10 | Refund unauthorized or incorrect charges | [Wrong charges, failed payments and plan provisioning](https://feedbackbench.com/criteria/account.billing_errors.md) | 16 | 22 |
| 11 | Restore original other-models allowance pool | [Price, allowance or plan terms changed](https://feedbackbench.com/criteria/limits.allowance_change.md) | 15 | 30 |
| 12 | Usage reset for new model or feature launch | [Quota reset timing and bonus or banked resets](https://feedbackbench.com/criteria/limits.reset_schedule.md) | 15 | 16 |

### 1. Native Android app

- Cursor, 2026-09-27, @cursor_ai (X): “cursor, where is the android version? 🤔 why is it still not available? iphone has it, and android users are still waiting... @cursor_ai 📱👀” [source](https://twitter.com/2000581649906667521/status/2104076132454985863)
- Cursor, 2026-09-26, @cursor_ai (X): “@cursor_ai please publish an android app. you have infinite tokens to spend on it.” [source](https://twitter.com/1051957462650314752/status/2103778855446011948)
- Cursor, 2026-09-24, @cursor_ai (X): “@cursor_ai @cursor_ai its been nearly 3 months since the ios app release. when can we expect to see the android app? we have been waiting patiently.” [source](https://twitter.com/1074060480199647232/status/2103039827545555360)

### 2. Release Composer 3 model

- Cursor, 2026-09-26, @cursor_ai (X): “composer 3? time to get back in the game @cursor_ai <strict_link>” [source](https://twitter.com/16070716/status/2103692283694739585)
- Cursor, 2026-09-23, r/cursor (Reddit): “grok is crazy expensive. but limit is kinda good. cursor need composer 3 and grok 5.0 to match claude / chatgpt” [source](https://www.reddit.com/r/cursor/comments/1wo1auh/gpt6_solluna_are_absolutely_cracked_and_busted/pbjbfuj/)
- Cursor, 2026-09-23, r/cursor (Reddit): “sol 6 and luna 6 is crazy efficient. sure they are not the "top-end" models, but they are still so good for most tasks and also so cheap comparable. i loved composer in the past for its kind of "efficiency" but yeah... i really hope we get something like composer 3 which can atleast a bit compete with those again.” [source](https://www.reddit.com/r/cursor/comments/1wo1auh/gpt6_solluna_are_absolutely_cracked_and_busted/pbj9crh/)

### 3. One-off usage limit reset now

- Cursor, 2026-09-25, @cursor_ai (X): “.@cursor_ai please for the love of god reset my usage i can't live without cursor for 20 days. i'm begging you, i'll be better this time.” [source](https://twitter.com/1701990401329033217/status/2103477234015281356)
- Cursor, 2026-09-24, @cursor_ai (X): “@cursor_ai ☕️😏 now reset usage so i can put my bots 🤖 back to work. 🤣😂😆😱💀👻🪦” [source](https://twitter.com/909198758/status/2102990399208087935)
- Cursor, 2026-09-22, @cursor_ai (X): “@cursor_ai i wait for a day where cursor will give a reset to everyone.” [source](https://twitter.com/1571938295625531394/status/2102489232895840332)

### 4. Unified subscription across linked products

- Cursor, 2026-09-26, @cursor_ai (X): “months and no action on unifying cursor/grok plans @spacexai? @cursor_ai has been cooking, but my ai budget is for two providers, and that's @anthropicai and cursor which means grok build gets fully cut out of the mix. don't tell me to use it for 4.7 when you make it so i cant” [source](https://twitter.com/995626692/status/2103901849485287638)
- Cursor, 2026-09-23, @cursor_ai (X): “the relationship between @cursor_ai @x @grok and @bot is very odd. you can subscribe to each separately, but each share benefits across each other. it’s very strange and not straightforward. i’d love to see @spacexai simplify this so it’s easier to know what to subscribe to based on our specific use case.” [source](https://twitter.com/1788271564649029632/status/2102870864266240305)
- Cursor, 2026-09-21, @cursor_ai (X): “migration complete. i’ve mirrored my claude and claude code workflows into grok, grok bot, and cursor. same operating system. new runtime. the workflows survived the move. what i’m waiting on now is simple: one integrated subscription across @grok, @bot , and @cursor_ai from @elonmusk and the @xai team. not three billing lines. one stack, full power, less friction. if the tools are meant to work together, the subscription should too.” [source](https://twitter.com/1396032511/status/2101849484279816480)

### 5. Add GPT-6 Astra model

- Cursor, 2026-09-11, @cursor_ai (X): “why hasn't @cursor_ai provided gpt-6 astra yet? wasn't it said that the cooperation would only stop in november?” [source](https://twitter.com/529069413/status/2098258235233100082)
- Cursor, 2026-09-09, @cursor_ai (X): “@cursor_ai where is gpt-6 astra?! been waiting for it every day!” [source](https://twitter.com/2446267116/status/2097832628032602188)
- Cursor, 2026-09-09, @cursor_ai (X): “@cursor_ai @grok do you know when gpt-6 astra will be available on cursor?” [source](https://twitter.com/1503310445566136321/status/2097586466008310208)

### 6. Add Grok 4.7 model

- Cursor, 2026-09-27, @cursor_ai (X): “yes.. elon and @cursor_ai. that’s your benchmark starring at you. make it work with grok <strict_link>” [source](https://twitter.com/1986276449280577536/status/2104206525019673025)
- Cursor, 2026-09-27, @cursor_ai (X): “@cursor_ai i have the cursor pro+ plan, but why am i not able to use the grok 4.7 model in it? <strict_link>” [source](https://twitter.com/2510853804/status/2104092513271509421)
- Cursor, 2026-09-23, r/cursor (Reddit): “grok is crazy expensive. but limit is kinda good. cursor need composer 3 and grok 5.0 to match claude / chatgpt” [source](https://www.reddit.com/r/cursor/comments/1wo1auh/gpt6_solluna_are_absolutely_cracked_and_busted/pbjbfuj/)

### 7. Higher overall usage limits

- Cursor, 2026-09-22, @cursor_ai (X): “this a problem, i agree. however it’s a economic problem. either extend the token limit or lower the token price and get 20x plan. i’m also a supergrok heavy user; however the benefits for €300 pm aren’t extraordinary vs other frontier models and apps. grok 4.7 isn’t mind blowing either. so grok bot does add a lot of value.” [source](https://twitter.com/1579852785063002112/status/2102520109839307150)
- Cursor, 2026-09-22, @cursor_ai (X): “@spacexai @cursor_ai give cursor more power so we don’t have to babysit usage mid-work. why can’t people have unlimited usage? we already pay a lot for a limited plan that dies halfway through a session. then wait. or spend more..? that’s not access. that’s a bottleneck. not asking for free.. asking for a fair price that lets you finish. give grok in cursor enough headroom that builders stay in flow instead of staring at a meter.” [source](https://twitter.com/497526709/status/2102508060811903191)
- Cursor, 2026-09-22, @cursor_ai (X): “@tesla @bot give more to x @premium+ subscribers, please! this and @cursor_ai would be 🧁” [source](https://twitter.com/17681343/status/2102435848620749114)

### 8. No silent model switching or downgrades

- Cursor, 2026-09-26, @cursor_ai (X): “@cursor_ai could you please stop switching my sessions to grok4.7? i know you want to push your new model, but it's not what i want an super-annoying. #customerfirst” [source](https://twitter.com/2717214655/status/2103820104416858532)
- Cursor, 2026-09-26, @cursor_ai (X): “no longer going to update cursor @cursor_ai @spacexai since you guys want to keep turning models on, and ignoring my settings, with them off. <strict_link>” [source](https://twitter.com/1437891983117279233/status/2103664293258633261)
- Cursor, 2026-09-25, r/cursor (Reddit): “yeah, i know about this popup, but i never accept is. additionally i've uploaded video and attached to the thread where you can see one case where the model changes on itself. i hope we can move the discussion from "it's your mistake" to "cursor changes models without permission"” [source](https://www.reddit.com/r/cursor/comments/1wpcg0f/i_just_lost_200_usd_because_cursor_switched_my/pbzzuo1/)

### 9. Additional or recurring bonus usage resets

- Cursor, 2026-09-26, @cursor_ai (X): “hey! can we get a cursor model reset please???? i won’t make it 2 more weeks! @cursor_ai @spacexai <strict_link>” [source](https://twitter.com/1635045021400580097/status/2103679538966446109)
- Cursor, 2026-09-22, @cursor_ai (X): “@cursor_ai @cursor_ai does this mean you'll give us usage refresh to celebrate?” [source](https://twitter.com/2098776405697839104/status/2102507070746763607)
- Cursor, 2026-09-22, @cursor_ai (X): “ok... @cursor_ai &amp; @xai y'all gonna watch your competitors give their users "resets" and ...not offer them on cursor plans? 🤔🤷♂️ <strict_link>” [source](https://twitter.com/1309409339824840704/status/2102453767698301170)

### 10. Refund unauthorized or incorrect charges

- Cursor, 2026-09-26, @cursor_ai (X): “i did not authorize any of these plans. it's been 8 days of absolute silence from @cursor_ai. this is a massive security &amp; billing flaw on your end. if this is not manually reviewed and refunded immediately, my next step is a formal fraud chargeback via stripe. (3/3) <strict_link>” [source](https://twitter.com/1513858901070200836/status/2103705967808651658)
- Cursor, 2026-09-25, @cursor_ai (X): “hey @cursor_ai, @poteto can someone help with a billing issue? i got a free month of pro+, never added a payment method, then got a $50 invoice after renewal. i tried to cancel but the unpaid invoice blocks it. i’ve asked for human review. please help void it and cancel renewal.” [source](https://twitter.com/1577735848283471874/status/2103515863462949272)
- Cursor, 2026-09-25, @cursor_ai (X): “@cursor_ai it’s been several days since i reported an unauthorized $236 ultra charge. i’ve followed up by email but haven’t received any update. please look into my ticket and help with the refund.” [source](https://twitter.com/2102638135620644865/status/2103357783156641863)

### 11. Restore original other-models allowance pool

- Cursor, 2026-09-25, @cursor_ai (X): “@fatih @cursor_ai would be cool to use if cursor didn’t reduced by -80% ultra allowance last month” [source](https://twitter.com/1878902328075366400/status/2103524408627220525)
- Cursor, 2026-09-24, @cursor_ai (X): “@cursor_ai a 7% token-cost cut is good engineering. pair it with honest ultra quotas. i'm on ultra via supergrok heavy. mid-subscription other models ~$400 → $100 — cursor confirmed heavy-linked ultra deliberately gets the smaller allowance. efficiency gains mean little if the entitlement was cut silently mid-cycle.” [source](https://twitter.com/1822950236593192961/status/2102928249193947286)
- Cursor, 2026-09-23, @cursor_ai (X): “@cursor_ai will you guys roll back to $400 for ultra credits?” [source](https://twitter.com/4556876521/status/2102797223021150671)

### 12. Usage reset for new model or feature launch

- Cursor, 2026-09-22, @cursor_ai (X): “@spacexai so don't we get a cursor reset like previous times when a new model drops? seeing everywhere some discount on grok 4.7 usage other than on cursor, so weird! @jediahkatz @anysph @cursor_ai @mntruell @ericzakariasson” [source](https://twitter.com/1464474213788569604/status/2102374857996669423)
- Cursor, 2026-09-21, @cursor_ai (X): “@bzagrodzki @cursor_ai would be nice to get a reset, when new model lands. <strict_link>” [source](https://twitter.com/351284352/status/2102124127809048661)
- Cursor, 2026-09-21, @cursor_ai (X): “hey @cursor_ai can we figure something out so i can test grok 4.7? 🥺 <strict_link>” [source](https://twitter.com/948560423355330560/status/2102098791298105584)

## Facts

| Fact | Value |
|---|---|
| Version | Composer 2.5 (agent model); Cursor 2.0 (multi-agent editor, git-worktree parallel agents) |
| Released | Composer 2.5: 2026-05-18. Composer 2: 2026-03-18/19. Cursor 2.0: 2025-10 |
| Price | Free (Hobby); Pro ~$20/mo credit-pool; Business/Enterprise custom. Composer 2.5 API: $0.50/$2.50 per 1M tokens standard, $3.00/$15.00 Fast (default) |
| Model | Composer 2.5 (proprietary) plus BYO access to Claude, GPT, Gemini |
| Surface | IDE (VS Code fork) |

## Sources

| Channel | Source | Posts |
|---|---|---|
| X | @cursor_ai | 9703 |
| Reddit | r/cursor | 8303 |
| Reddit | Posts that name it | 838 |
| Trustpilot | Trustpilot | 29 |
| G2 | G2 | 5 |

## Better than peers on

How much use a plan's price buys, IDE and editor integration, Which models are offered on a plan and when, Can do the user's kind of task, Plan-before-edit mode

## Worse than peers on

Quota reset timing and bonus or banked resets, Usage meter visibility and accuracy, Pay-as-you-go overage, fallback billing and spend caps, Using an existing subscription across tools, Connecting own API keys, local models and custom endpoints, Automatic model routing and fallback, Quality got worse or better over time, How the interface shows work, and what the user can configure, Mobile, remote-control and voice access, Support, refunds and issue handling

## All 63 criteria

Criterion love: 0.5 is the category norm. n: rated author-weeks.

### Paying and limits: Typical (customer love 0.506, n 2079)

| Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|
| [How much use a plan's price buys](https://feedbackbench.com/criteria/limits.plan_value.md) | Better than peers | 0.537 | 0.513–0.562 | 1001 | 427 | 574 |
| [Single prompt, model or effort level consumes disproportionate quota](https://feedbackbench.com/criteria/limits.burn_rate.md) | Typical | 0.525 | 0.487–0.558 | 621 | 119 | 502 |
| [Pricing and plan terms stated clearly and consistently](https://feedbackbench.com/criteria/billing.pricing_clarity.md) | Typical | 0.461 | 0.400–0.518 | 189 | 11 | 178 |
| [Price, allowance or plan terms changed](https://feedbackbench.com/criteria/limits.allowance_change.md) | Typical | 0.453 | 0.397–0.510 | 170 | 10 | 160 |
| [Quota reset timing and bonus or banked resets](https://feedbackbench.com/criteria/limits.reset_schedule.md) | Worse than peers | 0.442 | 0.399–0.481 | 136 | 15 | 121 |
| [Usage meter visibility and accuracy](https://feedbackbench.com/criteria/limits.usage_meter.md) | Worse than peers | 0.449 | 0.400–0.494 | 132 | 8 | 124 |
| [Pay-as-you-go overage, fallback billing and spend caps](https://feedbackbench.com/criteria/billing.overage_charges.md) | Worse than peers | 0.430 | 0.393–0.479 | 105 | 3 | 102 |
| [Using an existing subscription across tools](https://feedbackbench.com/criteria/billing.subscription_portability.md) | Worse than peers | 0.429 | 0.403–0.454 | 69 | 7 | 62 |
| [Short rolling usage window blocks or interrupts work](https://feedbackbench.com/criteria/limits.window_interrupts_work.md) | Typical | 0.540 | 0.496–0.581 | 51 | 13 | 38 |
| [Free tier and free model availability and limits](https://feedbackbench.com/criteria/billing.free_tier.md) | Typical | 0.482 | 0.457–0.509 | 43 | 20 | 23 |
| [Prompt cache hits, misses and invalidation](https://feedbackbench.com/criteria/limits.prompt_cache.md) | Too few posts | 0.501 | 0.485–0.518 | 22 | 11 | 11 |

Most recent posts:

- Praise, 2026-09-27, r/cursor (Reddit): “yes, but some have cursor anyway so its still "free bonus", even if not much” [source](https://www.reddit.com/r/cursor/comments/1wrbmrn/is_it_a_waste_of_money_to_run_opus_55_inside_of/pcc155q/)
- Praise, 2026-09-27, r/cursor (Reddit): “if you have the right setup, you can do almost all of your work in composer, which will last a long, long time. you can do this even for complicated tasks. i described it here: <strict_link>” [source](https://www.reddit.com/r/cursor/comments/1wrkzi5/am_i_cooked/pcdc7mq/)
- Praise, 2026-09-27, r/cursor (Reddit): “hmmm cursors limited pretty generous...” [source](https://www.reddit.com/r/cursor/comments/1wrkzi5/am_i_cooked/pcdewcc/)
- Complaint, 2026-09-27, r/cursor (Reddit): “i have given up trying .i can not afford it” [source](https://www.reddit.com/r/cursor/comments/1wjy8nh/is_there_any_way_to_make_cursor_grok_46_speak/pca4swh/)
- Complaint, 2026-09-27, r/cursor (Reddit): “opus in cursor absolutely ate through my ultra in a few hours unfortunately.” [source](https://www.reddit.com/r/cursor/comments/1wrbmrn/is_it_a_waste_of_money_to_run_opus_55_inside_of/pcbaloh/)
- Complaint, 2026-09-27, r/cursor (Reddit): “well no we don't. but if you did your job 2 years ago, why are you adding so much ai to it? i think about the project a lot before moving on to execution. i guess i dont have agents doing the thinking for me and going "/goal... !!!" but i really want to know how people spend $500 on ai every month.” [source](https://www.reddit.com/r/cursor/comments/1wr1ta7/i_feel_like_i_got_ripped_off/pcbbtdy/)

### Setting up and connecting: Typical (customer love 0.498, n 431)

| Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|
| [MCP servers, plugins, skills and hooks](https://feedbackbench.com/criteria/setup.extensions_mcp.md) | Typical | 0.481 | 0.447–0.515 | 123 | 55 | 68 |
| [IDE and editor integration](https://feedbackbench.com/criteria/setup.ide_integration.md) | Better than peers | 0.576 | 0.547–0.604 | 111 | 74 | 37 |
| [Install, launch and sign-in](https://feedbackbench.com/criteria/setup.install_signin.md) | Typical | 0.479 | 0.441–0.515 | 88 | 13 | 75 |
| [Connecting own API keys, local models and custom endpoints](https://feedbackbench.com/criteria/setup.provider_byok_local.md) | Worse than peers | 0.431 | 0.401–0.460 | 86 | 26 | 60 |
| [Onboarding, discoverability and documentation](https://feedbackbench.com/criteria/setup.onboarding_docs.md) | Typical | 0.500 | 0.468–0.531 | 42 | 8 | 34 |

Most recent posts:

- Praise, 2026-09-27, r/cursor (Reddit): “so here's my issue with grok models up until 3.x\~ they were training their own models on their own data. in house model, in house data. spacex sees what cursor is doing, which is basically using all the enterprise data flow and their retail use towards training their own in-house model and they want in on it too. here's the kicker, they all start with the same base, it's all kimi 2.5 under the ho…” [source](https://www.reddit.com/r/cursor/comments/1wrktgc/am_i_the_only_one_who_thinks_grok_47_is_actually/pcdbwh3/)
- Praise, 2026-09-27, r/cursor (Reddit): “move to the desktop app, i left ide a month ago and will never look back” [source](https://www.reddit.com/r/cursor/comments/1wrbmrn/is_it_a_waste_of_money_to_run_opus_55_inside_of/pcfbz09/)
- Praise, 2026-09-27, r/cursor (Reddit): “coding in windows is just joking with the filesystem . as cursor is based on vscode, at least udñse remote developement with wsl wich is highly integrated and 1 click install” [source](https://www.reddit.com/r/cursor/comments/1wq0bdq/cursor_wiped_out_a_guys_entire_drive/pcfz2m8/)
- Complaint, 2026-09-27, r/cursor (Reddit): “i feel you. i went down this exact rabbit hole a few months ago — loved the tile view and the composer workflow, but the moment i hit a real refactor session the quota wall killed the flow. short answer: cursor doesn't let you bring your own key for their $20 plan, and the byok workaround people hack together is brittle (you lose the indexing, the diff view, the agent loops). what actually solved…” [source](https://www.reddit.com/r/cursor/comments/1wqq46y/anyone_managed_to_use_cursors_ui_with_their/pcd5l9b/)
- Complaint, 2026-09-27, @cursor_ai (X): “cursor, where is the android version? 🤔 why is it still not available? iphone has it, and android users are still waiting... @cursor_ai 📱👀” [source](https://twitter.com/2000581649906667521/status/2104076132454985863)
- Complaint, 2026-09-27, @cursor_ai (X): “@sicaniansun @cursor_ai it's been so long and there's still no android version released.” [source](https://twitter.com/2000581649906667521/status/2104152763622101005)

### Choosing models: Typical (customer love 0.489, n 959)

| Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|
| [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | Better than peers | 0.550 | 0.515–0.582 | 361 | 128 | 233 |
| [Automatic model routing and fallback](https://feedbackbench.com/criteria/models.routing_auto.md) | Worse than peers | 0.467 | 0.433–0.499 | 343 | 73 | 270 |
| [Quality got worse or better over time](https://feedbackbench.com/criteria/models.quality_drift.md) | Worse than peers | 0.443 | 0.404–0.483 | 314 | 54 | 260 |
| [Reasoning effort setting and its defaults](https://feedbackbench.com/criteria/models.effort_control.md) | Too few posts | 0.512 | 0.492–0.535 | 29 | 16 | 13 |

Most recent posts:

- Praise, 2026-09-27, r/cursor (Reddit): “sonnet's also faster, we don't need opus/fable level intelligence for most tasks 🤷🏿♂️” [source](https://www.reddit.com/r/cursor/comments/1wp6j44/is_this_supposed_to_be_good_news/pcbgqaz/)
- Praise, 2026-09-27, r/cursor (Reddit): “cursor is the complete package. powerful ide, multiple models. generous composer and grok. you have grok bot too and environment vm.” [source](https://www.reddit.com/r/cursor/comments/1wrkzi5/am_i_cooked/pcdgc47/)
- Praise, 2026-09-27, r/cursor (Reddit): “this actually feels reasonable! i also feel like we don't need too much intelligence for most of the task and grok is a goof starting point” [source](https://www.reddit.com/r/cursor/comments/1wp6j44/is_this_supposed_to_be_good_news/pcej6hc/)
- Complaint, 2026-09-27, r/cursor (Reddit): “i think in auto if it gets routed to an expensive model , it will cost you the full pricing of the model since they removed the lower/discounted pricing from it. previously auto only used composer or grok” [source](https://www.reddit.com/r/cursor/comments/1wrddz6/why_is_the_other_models_usage_is_being_used_so/pcbz9uf/)
- Complaint, 2026-09-27, r/cursor (Reddit): “actually my both ugage are over and my task is specific to grok 4.6 i just need that model... (other model in capable of doing that). or open source model.” [source](https://www.reddit.com/r/cursor/comments/1wrkzi5/am_i_cooked/pcdcxno/)
- Complaint, 2026-09-27, r/cursor (Reddit): “on the other models / which settings ask, auto is what chewed through mine when it routed into an expensive model at full list price, so i pin grok 4.6 now and on a bad stretch other models still jumped maybe \~40% in under an hour” [source](https://www.reddit.com/r/cursor/comments/1wrddz6/why_is_the_other_models_usage_is_being_used_so/pcdioo5/)

### Instructing and context: Better than peers (customer love 0.542, n 318)

| Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|
| [Memory and state carried across sessions](https://feedbackbench.com/criteria/context.session_memory.md) | Typical | 0.524 | 0.494–0.556 | 110 | 64 | 46 |
| [Finding the right files in the codebase](https://feedbackbench.com/criteria/context.codebase_retrieval.md) | Typical | 0.523 | 0.496–0.551 | 64 | 36 | 28 |
| [Persistent project rules files are read and obeyed](https://feedbackbench.com/criteria/context.instruction_files.md) | Typical | 0.487 | 0.457–0.515 | 62 | 29 | 33 |
| [Direct in-prompt instructions and caps are followed](https://feedbackbench.com/criteria/context.instruction_following.md) | Typical | 0.479 | 0.451–0.505 | 50 | 10 | 40 |
| [Output degrades as the context window fills](https://feedbackbench.com/criteria/context.long_context_decay.md) | Too few posts | 0.521 | 0.489–0.552 | 28 | 7 | 21 |
| [Context compaction keeps what matters, cheaply and quickly](https://feedbackbench.com/criteria/context.compaction.md) | Too few posts | 0.507 | 0.488–0.527 | 21 | 9 | 12 |
| [Asks the user versus guessing](https://feedbackbench.com/criteria/context.clarifying_questions.md) | Too few posts | 0.491 | 0.478–0.504 | 10 | 2 | 8 |
| [Images, PDFs and file attachments as input](https://feedbackbench.com/criteria/context.attachments.md) | Too few posts | 0.496 | 0.484–0.510 | 7 | 2 | 5 |

Most recent posts:

- Praise, 2026-09-27, r/cursor (Reddit): “i hear you. i often need to pause it if i'm in a flow state but for day to day stuff i really like it. it seems to get better the larger the code base and the more it can identify patterns. but that my also be my confirmation bias. i'm trying to use less agentic methods so i force myself to know what's going on an the autocomplete is a nice balance for me.” [source](https://www.reddit.com/r/cursor/comments/1wquwo9/autocomplete/pc9th3g/)
- Praise, 2026-09-27, r/cursor (Reddit): “i find that it’s actually pretty good at understanding a codebase accurately. sometimes i feel like opus will take a shortcut or get sidetracked. however, once the understanding is there, opus is a better planner and executor.” [source](https://www.reddit.com/r/cursor/comments/1wrktgc/am_i_the_only_one_who_thinks_grok_47_is_actually/pcdhko1/)
- Praise, 2026-09-27, r/cursor (Reddit): “composer does exactly what you ask it to do even if it takes a few prompts to finish grok will do it all and add 10 things i didn't ask for so i tell it i didn't ask for those things and it says 'you're right i'm so sorry' then it adds 2 other things i didn't want or it will change something that breaks everything. so you ask it to fix it. oh, so sorry, here's 2 more things you didn't ask for.” [source](https://www.reddit.com/r/cursor/comments/1wrktgc/am_i_the_only_one_who_thinks_grok_47_is_actually/pcdlxql/)
- Complaint, 2026-09-27, @cursor_ai (X): “@cursor_ai should reconsider cx of follow up questions after execution of approved plan started. it hanged entire authonomy.” [source](https://twitter.com/255140211/status/2104192810089943545)
- Complaint, 2026-09-27, r/webdev (Reddit): “the en dash in pikspec is doing a lot of work there, your store url has %e2%80%93 sitting right in the middle of it so every link you ever paste looks like it went through a redirector. good luck with that one. [design.md](<strict_link>) is the part i'd actually use, though cursor ignores it unless i @ it in every single message.” [source](https://www.reddit.com/r/webdev/comments/1wrriz4/made_a_chrome_extension_so_cursorclaude_stop/pcfb2p6/)
- Complaint, 2026-09-27, r/AI_Agents (Reddit): “i'd want the memory part to actually work across different projects without me having to explain the same thing twice. my current setup with cursor is basically just me telling it the same architecture rules over and over every time i start a new chat the background agents thing could be useful too but i think the real test is whether the swarm mode produces anything coherent or just burns through…” [source](https://www.reddit.com/r/AI_Agents/comments/1wrv4vw/im_building_ai_swarms_that_research_debate_and/pcg4p26/)

### Doing the work: Better than peers (customer love 0.533, n 1252)

| Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|
| [Can do the user's kind of task](https://feedbackbench.com/criteria/work.capability.md) | Better than peers | 0.555 | 0.526–0.584 | 639 | 444 | 195 |
| [Subagents, parallel agents and orchestrators](https://feedbackbench.com/criteria/work.multi_agent_orchestration.md) | Typical | 0.535 | 0.499–0.571 | 211 | 143 | 68 |
| [Risky or irreversible actions without confirmation](https://feedbackbench.com/criteria/work.destructive_actions.md) | Typical | 0.475 | 0.445–0.509 | 61 | 7 | 54 |
| [Long unattended runs and goal/loop mode](https://feedbackbench.com/criteria/work.long_running_autonomy.md) | Typical | 0.503 | 0.472–0.534 | 56 | 44 | 12 |
| [Spins, loops or gets stuck without progress](https://feedbackbench.com/criteria/work.stuck_loops.md) | Typical | 0.459 | 0.428–0.503 | 54 | 1 | 53 |
| [Frontend and visual UI output](https://feedbackbench.com/criteria/work.frontend_ui.md) | Typical | 0.498 | 0.473–0.521 | 50 | 27 | 23 |
| [Breaks existing code or reintroduces bugs](https://feedbackbench.com/criteria/work.regressions_introduced.md) | Typical | 0.510 | 0.482–0.538 | 50 | 11 | 39 |
| [Does unrequested work or over-engineers](https://feedbackbench.com/criteria/work.scope_overreach.md) | Typical | 0.473 | 0.435–0.513 | 50 | 2 | 48 |
| [Git commits, branches and sync](https://feedbackbench.com/criteria/work.git_workflow.md) | Typical | 0.485 | 0.461–0.508 | 45 | 12 | 33 |
| [Computer use and browser control](https://feedbackbench.com/criteria/work.computer_browser_use.md) | Typical | 0.482 | 0.458–0.505 | 38 | 17 | 21 |
| [Tool approval prompts and autonomy modes](https://feedbackbench.com/criteria/work.permission_prompts.md) | Typical | 0.505 | 0.475–0.532 | 37 | 10 | 27 |
| [Plan-before-edit mode](https://feedbackbench.com/criteria/work.plan_mode.md) | Better than peers | 0.524 | 0.501–0.547 | 37 | 23 | 14 |
| [Safety filters block legitimate coding tasks](https://feedbackbench.com/criteria/work.safety_refusals.md) | Typical | 0.484 | 0.456–0.515 | 30 | 2 | 28 |
| [Length and clarity of replies, summaries and comments](https://feedbackbench.com/criteria/work.response_verbosity.md) | Too few posts | 0.485 | 0.466–0.509 | 22 | 3 | 19 |
| [Diagnosing and fixing reported bugs](https://feedbackbench.com/criteria/work.bug_diagnosis.md) | Too few posts | 0.502 | 0.485–0.520 | 18 | 13 | 5 |
| [Stops mid-task or answers instead of acting](https://feedbackbench.com/criteria/work.premature_stop.md) | Too few posts | 0.487 | 0.478–0.495 | 8 | 0 | 8 |
| [Caves to or argues with the user's judgement](https://feedbackbench.com/criteria/work.sycophancy_pushback.md) | Too few posts | 0.498 | 0.489–0.517 | 5 | 1 | 4 |
| [Games checks instead of fixing the problem](https://feedbackbench.com/criteria/work.reward_hacking.md) | Too few posts | 0.495 | 0.490–0.499 | 4 | 0 | 4 |

Most recent posts:

- Praise, 2026-09-27, r/cursor (Reddit): “right there with you - not an elon fan, but have had grok 4.6 and now 4.7 as my everyday driver since composer 2.5. all entirely competent models.” [source](https://www.reddit.com/r/cursor/comments/1wrktgc/am_i_the_only_one_who_thinks_grok_47_is_actually/pcdcwpu/)
- Praise, 2026-09-27, r/cursor (Reddit): “yep 👍🏼. actually i used agent i needed that actually to run autonomous task till 2days constant...” [source](https://www.reddit.com/r/cursor/comments/1wrkzi5/am_i_cooked/pce76ru/)
- Praise, 2026-09-27, r/cursor (Reddit): “did some autonomous tasks...(2days constant workout with 8-9 agents running)” [source](https://www.reddit.com/r/cursor/comments/1wrkzi5/am_i_cooked/pce8fh2/)
- Complaint, 2026-09-27, r/cursor (Reddit): “cursor agents are fast at *retrying*. the expensive part for us was retrying the **same** fail — bad path, wrong toolchain, flake that already had a known fix in another session. we keep a small oss prior-art index (claimidx, apache-2.0) beside the agent: ask before grinding, apply + verify, then publish a compact claim. retrieved remedies are evidence for the model, not auto-executed patches. if…” [source](https://www.reddit.com/r/cursor/comments/1wqkwtj/your_cursor_plan_already_spins_cloud_agents_why/pc9t0l9/)
- Complaint, 2026-09-27, r/cursor (Reddit): “one-shotting is the issue. unless it's something as basic as a chrome extension, i never one-shotting.” [source](https://www.reddit.com/r/cursor/comments/1wqppuh/grok_is_shutting_down_apps_now/pcahmk0/)
- Complaint, 2026-09-27, r/cursor (Reddit): “cursor is crap compared to a pro plan on claude code.” [source](https://www.reddit.com/r/cursor/comments/1wrbmrn/is_it_a_waste_of_money_to_run_opus_55_inside_of/pcd0den/)

### Checking and finishing: Better than peers (customer love 0.542, n 155)

| Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|
| [Reviewing and approving the agent's changes](https://feedbackbench.com/criteria/verify.change_review_ui.md) | Typical | 0.526 | 0.498–0.552 | 52 | 26 | 26 |
| [Builds, tests or runs its own changes](https://feedbackbench.com/criteria/verify.self_testing.md) | Typical | 0.502 | 0.478–0.526 | 51 | 33 | 18 |
| [Agent-performed code review finds real issues](https://feedbackbench.com/criteria/verify.agent_code_review.md) | Typical | 0.515 | 0.491–0.542 | 36 | 30 | 6 |
| [Claims work is done or fixed when it is not](https://feedbackbench.com/criteria/verify.false_completion.md) | Too few posts | 0.509 | 0.473–0.557 | 19 | 2 | 17 |

Most recent posts:

- Praise, 2026-09-27, @cursor_ai (X): “@cursor_ai verifying the deploy, not just the diff, is such a smart way to close the loop. a monitoring plan written with the change is the step most teams skip. would love this for mobile releases too, where a regression lives until the next store review.” [source](https://twitter.com/1667644418768375808/status/2104295956967821398)
- Praise, 2026-09-26, r/cursor (Reddit): “i'd switch. if the ui is what's slowing you down, that costs you more than the quota ever will, and cursor's multi-chat and diff review are genuinely nicer. just don't cancel codex yet: run cursor on your real repo for a few heavy days and watch the usage meter. if you're burning it on tiny edits, that's the workflow leaking, not the plan.” [source](https://www.reddit.com/r/cursor/comments/1wptv9n/thinking_of_switching_from_codex_to_cursor/pc4r5px/)
- Praise, 2026-09-25, r/cursor (Reddit): “hello there. so regarding your question, from experience, i have been subscribed with cursor for around two years, a yearly subscription. the usage limit actually is the best you will ever get. i will share with you a photo from my usage, so you can see that if you use the composer 2.5, you get around 2 billion tokens from my current workload, which is as a full-time developer working on multiple…” [source](https://www.reddit.com/r/cursor/comments/1wptv9n/thinking_of_switching_from_codex_to_cursor/pbywqpt/)
- Complaint, 2026-09-27, @cursor_ai (X): “@cursor_ai the plan is the part that lies. ours greened on the deploy log while checkout 500'd. if the monitor doesn't hit the user path, it's just watching itself.” [source](https://twitter.com/1835841692852682752/status/2104186619410719018)
- Complaint, 2026-09-27, @cursor_ai (X): “@cjbell_ @cursor_ai agent branch commits hidden till pr is frustrating, i've hit that markdown plan viewer shuffle too” [source](https://twitter.com/184674873/status/2104338238085509151)
- Complaint, 2026-09-26, @cursor_ai (X): “@cursor_ai self-verification on cursorbench is a model skill, not a permission. an agent that checks its own diffs can still ship the wrong change if the only reviewer is the same loop that wrote it.” [source](https://twitter.com/1288646414394896389/status/2103764156361150570)

### Interface and sessions: Typical (customer love 0.494, n 560)

| Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|
| [Cloud and remote sandbox execution](https://feedbackbench.com/criteria/surfaces.cloud_sessions.md) | Typical | 0.503 | 0.472–0.536 | 207 | 138 | 69 |
| [How the interface shows work, and what the user can configure](https://feedbackbench.com/criteria/ui.display_settings.md) | Worse than peers | 0.451 | 0.418–0.486 | 201 | 53 | 148 |
| [Mobile, remote-control and voice access](https://feedbackbench.com/criteria/surfaces.remote_mobile.md) | Worse than peers | 0.437 | 0.408–0.468 | 120 | 40 | 80 |
| [Saving, switching, resuming and rewinding sessions](https://feedbackbench.com/criteria/ui.session_history.md) | Typical | 0.506 | 0.476–0.537 | 59 | 19 | 40 |
| [Stopping and steering a running agent](https://feedbackbench.com/criteria/ui.interrupt_steer.md) | Too few posts | 0.489 | 0.473–0.505 | 18 | 6 | 12 |

Most recent posts:

- Praise, 2026-09-27, @cursor_ai (X): “@chatgpt codex cloud is such a shit show. @cursor_ai cloud agent is miles ahead. not too sure about @claudeai though.” [source](https://twitter.com/1215028704/status/2104076818563379504)
- Praise, 2026-09-27, @cursor_ai (X): “this feels like a cheat code. recorded a skill via @claudeai desktop on my mac. created a new repo in @github to exercise the new skill. loaded the repo into @cursor_ai so i can answer some of the skill's questions on my phone while i wait for my son's soccer game.” [source](https://twitter.com/10245302/status/2104273772199239694)
- Praise, 2026-09-26, r/cursor (Reddit): “i haven’t found any ui that is as nice to use as cursor, but you can try zed. if you maximise the agent panel then the ui is good. you can then use claude code + opus 5.5 via acp in zed. for dictation you can either pay for wisprflow or use typewhisper for free” [source](https://www.reddit.com/r/cursor/comments/1wptv9n/thinking_of_switching_from_codex_to_cursor/pc3n7mi/)
- Complaint, 2026-09-27, @cursor_ai (X): “dear cursor team, please add feedback to your agents - might be because i suck with coding, but if an agent doesn't respond and doesnt give me visual feedback that is terrible ux because i won't know that it doesn't do what i want it to do until it's already done it @cursor_ai” [source](https://twitter.com/1848033478865993728/status/2104196348778266844)
- Complaint, 2026-09-27, r/vibecoding (Reddit): “yes i know, all bigger ai labs have cloud work now, but they still revolve around session management that you have to steer like in codex. sdlc is lacking, nor any redaction, deduplication, filter or teams functionality you have to build the whole setup yourself. i've seen some like cursor cloud now starting to add teamwork but still many features are lacking for real production use yet. best curr…” [source](https://www.reddit.com/r/vibecoding/comments/1wrm0ec/vibecoded_a_whole_software_factory_now_it_builds/pcec5hk/)
- Complaint, 2026-09-26, @cursor_ai (X): “@cursor_ai again feels like meh! first the laggy ide then the messed up layouts now the usage and models 3rd time in history i cancelled my subscription in less than a week for cursor. they did make a comeback most of the time but now codex just feels miles ahead” [source](https://twitter.com/1213841290825043969/status/2103697827318632623)

### Reliability and speed: Worse than peers (customer love 0.446, n 690)

| Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|
| [Outages, server errors and capacity or rate errors](https://feedbackbench.com/criteria/rel.service_errors.md) | Typical | 0.488 | 0.420–0.551 | 257 | 17 | 240 |
| [Latency, throughput and fast mode](https://feedbackbench.com/criteria/rel.response_speed.md) | Typical | 0.468 | 0.436–0.500 | 211 | 68 | 143 |
| [Client crashes, freezes and failed tool execution](https://feedbackbench.com/criteria/rel.client_failures.md) | Typical | 0.481 | 0.414–0.544 | 200 | 12 | 188 |
| [Updates break working setups](https://feedbackbench.com/criteria/rel.update_breakage.md) | Typical | 0.494 | 0.454–0.532 | 62 | 9 | 53 |

Most recent posts:

- Praise, 2026-09-27, r/cursor (Reddit): “you are right about part of it but wrong about fast being slower than normal . that could never happen. fast will have priority on resources. normal comes seconds . in rare cases normal could be same speed as fast . the speed of normal varies up and down. best case its same as fast . but fast almost always same speed for a certain window of time .” [source](https://www.reddit.com/r/cursor/comments/1wp2ped/i_compared_cursor_composer_25_normal_vs_fast/pccmxhn/)
- Praise, 2026-09-27, r/cursor (Reddit): “i’m running an m1 air w 16gb and it runs like a breeze. maybe it’s the actual workload you have it running?” [source](https://www.reddit.com/r/cursor/comments/1wmj0pw/does_cursor_make_anyone_elses_computer_extremely/pcdw03l/)
- Praise, 2026-09-27, r/cursor (Reddit): “sorry but this sounds like an operator problem. my system was running slow at one point so i prompted it to optimize my system. haven’t had a problem since, that was about 5 months ago.” [source](https://www.reddit.com/r/cursor/comments/1wqppuh/grok_is_shutting_down_apps_now/pcf5mxo/)
- Complaint, 2026-09-27, r/cursor (Reddit): “a few minutes/instant. switched to claude yesterday, fully operational on 3 very different projects.” [source](https://www.reddit.com/r/cursor/comments/1wq3m71/im_out/pc9w7tt/)
- Complaint, 2026-09-27, r/cursor (Reddit): “cursor hanging on taking longer than expected after the shell already finished is the same false busy lie as waiting for subagent. i kill that agent pane first and reopen the folder so the host actually resets. full app restart helps less than clearing the hung session. if the spinner comes back on the next prompt the host is sick not the model.” [source](https://www.reddit.com/r/cursor/comments/1wquvoq/taking_longer_than_expected/pccfywa/)
- Complaint, 2026-09-27, r/cursor (Reddit): “i don't feel similarly. much slower, way less accurate” [source](https://www.reddit.com/r/cursor/comments/1wrktgc/am_i_the_only_one_who_thinks_grok_47_is_actually/pcdnkab/)

### Account and support: Worse than peers (customer love 0.436, n 524)

| Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|
| [Support, refunds and issue handling](https://feedbackbench.com/criteria/account.support.md) | Worse than peers | 0.431 | 0.389–0.468 | 324 | 36 | 288 |
| [Wrong charges, failed payments and plan provisioning](https://feedbackbench.com/criteria/account.billing_errors.md) | Typical | 0.508 | 0.389–0.601 | 184 | 4 | 180 |
| [Account bans and access restrictions](https://feedbackbench.com/criteria/account.bans_restrictions.md) | Typical | 0.488 | 0.426–0.546 | 83 | 4 | 79 |
| [Data retention, training use and deployment isolation](https://feedbackbench.com/criteria/account.data_privacy.md) | Typical | 0.495 | 0.465–0.528 | 51 | 11 | 40 |

Most recent posts:

- Praise, 2026-09-25, r/cursor (Reddit): “"random invoices"... sure thing buddy, maybe you enabled on demand charging on your own. there is nothing random about cursor charging, but go to the greener grass by all means” [source](https://www.reddit.com/r/cursor/comments/1wph9rt/downgrading_from_200_to_20/pbxpxqr/)
- Praise, 2026-09-25, @cursor_ai (X): “@goldenberglior @github @cursor_ai my problem has been solved did you contact support? if so, i sent them another message on the same ticket. i don't know why, but they replied to the second message within a minute. <strict_link>” [source](https://twitter.com/1492254120530612230/status/2103519251910795649)
- Praise, 2026-09-24, @cursor_ai (X): “using muse in cursor right now, they reached out to me and wanted to sell me more than i can afford thank you @cursor_ai team some day soon i hope to get there where i can pay that! :) appreciate the credits!!” [source](https://twitter.com/1978977371584708608/status/2103179596174950500)
- Complaint, 2026-09-27, r/cursor (Reddit): “just dispute it with your bank. you’ll get nowhere with cursor support.” [source](https://www.reddit.com/r/cursor/comments/1wrewe3/disappointed_with_cursor_support_handling_an/pccf8id/)
- Complaint, 2026-09-27, r/cursor (Reddit): “i cancelled my cursor subscription months ago, and really realized how bad it was once i switched to cc. really frustrated to not have switched sooner. not to mention cursor censored my feedback on the support forum when i pointed out that there are too many bugs.” [source](https://www.reddit.com/r/cursor/comments/1wrbmrn/is_it_a_waste_of_money_to_run_opus_55_inside_of/pcd1lpo/)
- Complaint, 2026-09-27, r/cursor (Reddit): “i cancelled my cursor subscription in april and haven’t logged in since may. yesterday, my card was charged $200, and when i checked the account, i saw that cursor ultra had been activated. i did not activate it or approve the payment. when i first reported this, there was no usage showing. now i can see some usage appearing, but it is not mine. i don’t even have cursor installed on my computer no…” [source](https://www.reddit.com/r/cursor/comments/1wrewe3/disappointed_with_cursor_support_handling_an/pcd6dlt/)
