# Devin (Cognition)

Feedback Bench, coding agents, built 2026-10-01, window 2026-08-31 to 2026-09-27. Web page: https://feedbackbench.com/#/agent/devin

| Measure | Value |
|---|---|
| Rank | =5 of 17 (rank range 5–6) |
| Feedback Score | 55.9 (95% interval 55.1–56.5) |
| Popularity | 0.514 (share of voice 3.66%) |
| Customer love | 0.607 (95% interval 0.591–0.622) |
| Top quadrant | yes |
| Authors | 3603 |
| Posts counted | 6711 |
| Posts that judge the agent | 3237 |
| Criteria better / worse than peers | 9 / 2 of 63 |

## The brief

Written by Claude Opus 5.5 from 95 labelled posts and the numbers on this page. Interpretation, not measurement: every quote is verbatim and links to its post.

**A strong engineer in the cloud, throttled by its own meter.**

TL;DR:

- SWE-2 on the Pro plan is the deal users keep quoting, at least while the promo lasts.
- Rolling usage windows, slow responses and desktop crashes break flow more often than the model fails.
- Cloud sessions, model choice and long unattended runs are where Devin beats peers.

### What hurts

- **Usage windows cut sessions off mid-flow** ([Short rolling usage window blocks or interrupts work](https://feedbackbench.com/criteria/limits.window_interrupts_work.md), [Single prompt, model or effort level consumes disproportionate quota](https://feedbackbench.com/criteria/limits.burn_rate.md), [Usage meter visibility and accuracy](https://feedbackbench.com/criteria/limits.usage_meter.md)). Short rolling windows and per-session quota spikes stop work cold, so users who expect long runs get parked for half an hour instead.
  This is the one usage criterion where Devin lands worse than peers, and almost every post is a complaint. Users describe a 5-hour window draining in minutes, parallel tasks triggering lockouts, and a single heavy-model session eating a day's allowance plus half the week.
  
  The pain is unpredictability. Users say quotas feel hidden until they hit one, so a plan that looks generous on paper interrupts real work.
  Evidence:
  - Complaint, Devin, @cognition, 2026-09-11: “@doozieakshay @da7_tech @devindesktop @cognition not unlimited. heavy rate limiting. 2 parallel tasks = shut me down for 30 minutes. not happening for you?” [source](https://twitter.com/1430709146605740035/status/2098316309335359979)
  - Complaint, Devin, @DevinAI, 2026-09-05: “@ariskaa_ai @k2sbhai @devinai but the 5h limit is finishing too fast in just 5 to 10 minutes.” [source](https://twitter.com/1101398178350260230/status/2096282929794265333)
  - Complaint, Devin, @cognition, 2026-09-20: “@pluggsupply @dabit3 @devinai @cognition @da7_tech hidden 5-hour quotas that cut you off mid flow are annoying. i've got leftover qwen/k3 (bailian) on pay-per-use — key stays with me. dm if you want a soft quote.” [source](https://twitter.com/2092015317358989312/status/2101713090752454896)
  - Complaint, Devin, r/windsurf, 2026-09-06: “i'm posting to know wheter this is a bug or is it because i just bought it and its reseting on sunday.... a single astra session which did not complete fully immediately by sending the message it filled my daily and half of the week, is this normal? or first day shennaningans? <strict_link>” [source](https://www.reddit.com/r/windsurf/comments/1w8i4r6/one_gpt_astra_session_took_half_of_my_weekly/)

- **SWE-2 is slow under load** ([Latency, throughput and fast mode](https://feedbackbench.com/criteria/rel.response_speed.md)). Response latency is worse than peers, and users blame server capacity and slow code search rather than the model's reasoning.
  Users like what SWE-2 produces but not how long it takes. Posts describe long waits between prompts, slow file searching, and fast context not kicking in. Some call it slower than SWE-1.7 even though it needs fewer steps.
  
  Speed reports vary by surface and region. One user says cloud now runs faster than the desktop app, and another in Latin America reports no lag. That points to capacity, not the model itself.
  Evidence:
  - Complaint, Devin, @cognition, 2026-09-18: “@fei2411 @cognition @cognition please expand the capacity of your devin servers to boost the speed of swe-2. it's currently too slow.” [source](https://twitter.com/1159835302275346433/status/2100823113516941327)
  - Complaint, Devin, r/windsurf, 2026-09-15: “compared to swe 1.7, it's slow. it does work well within less steps compared to swe 1.7 and also thinks less in terms of tokens. code surfing, searching also seems to take a lot of time, it also doesn't trigger fast context. i don't know if the model itself is consuming their bandwidth or something else.” [source](https://www.reddit.com/r/windsurf/comments/1wd432p/swe2_first_experiences/p9ymvom/)
  - Complaint, Devin, @DevinAI, 2026-09-23: “@hraness @devinai @zeddotdev i send one prompt, then spend 20 minutes contemplating life, then the next prompt...” [source](https://twitter.com/14373276/status/2102810919101460606)
  - Praise, Devin, @cognition, 2026-09-19: “devin swe-2 from 75% discount to 100% free in cloud is a big win for me on pro plan. it's notably faster now on cloud then it was on desktop app and it's intelligence is impressive for what it's cost! @devindesktop @cognition @scottwu46 #devin #aigen” [source](https://twitter.com/1765566409168261120/status/2101186910752473098)

- **Desktop crashes and SWE-2 freezes** ([Client crashes, freezes and failed tool execution](https://feedbackbench.com/criteria/rel.client_failures.md), [Updates break working setups](https://feedbackbench.com/criteria/rel.update_breakage.md)). Client failures draw almost no praise, as users report repeated desktop crashes, hour-long freezes in planning mode, and runaway disk usage after sessions.
  Nearly every post on this criterion is a complaint. Users describe Devin Desktop crashing again and again, SWE-2 hanging on a large planning command until they switch back to an older model, and one user finding tens of gigabytes of disk consumed after sessions.
  
  A few users note that the agent recovers cleanly from internal tool errors. Recovery inside a run does not help when the client itself goes down.
  Evidence:
  - Complaint, Devin, @cognition, 2026-09-14: “@cognition swe-2 freeze completely when using /megaplan, after wait for an hour, got no result, switch to swe 1.7 lightning max everything works well <strict_link>” [source](https://twitter.com/2004023993314213889/status/2099484210235105621)
  - Complaint, Devin, @cognition, 2026-09-16: “@cognition why is devin eating 49gb of my disk space after running sessions?” [source](https://twitter.com/1485561275828641799/status/2100259383422652423)
  - Complaint, Devin, @cognition, 2026-09-18: “hmm. idk why devin desktop keeps on crashing @devindesktop @cognition” [source](https://twitter.com/1675486335942168578/status/2100800706005684528)
  - Praise, Devin, r/windsurf, 2026-09-08: “just thought i'd post this as an example of recovery; i'm using glm. devin errored, i didn't check the details, but it looked worth flagging so i stopped what glm was doing, which i always do if things are going not where i think they should. i flagged up the issue and it handled it well, which is what it generally seems to do. other ide's probably do too, but i appreciate how things don't fall apart. \--- **there was a json error when you were working (an internal error presumably from a missing status field)** *yes, i hit a transient json validation error on one todo\_write call (a missing status field) — i corrected it on the next call and the list is now tracking properly. continuing with the implementation.*” [source](https://www.reddit.com/r/windsurf/comments/1wasbyy/devin_recovery_example/)

- **Free tier terms keep shifting** ([Free tier and free model availability and limits](https://feedbackbench.com/criteria/billing.free_tier.md), [Pricing and plan terms stated clearly and consistently](https://feedbackbench.com/criteria/billing.pricing_clarity.md), [Price, allowance or plan terms changed](https://feedbackbench.com/criteria/limits.allowance_change.md)). Free models that users built habits on got quietly rate-limited or pulled, and stated pricing rarely matches what the meter enforces.
  Pricing clarity is close to all complaints. Users loved a free GLM tier, then report new rate limits and a daily quota after a CLI update, on a model still described as free. Others say a free trial ran out before they sent a prompt.
  
  The complaint is not the price. It is the gap between what the plan says and what the meter does. Published exact limits per plan is a standing request.
  Evidence:
  - Complaint, Devin, @cognition, 2026-09-18: “@dabit3 i'm disappointed with the free version of @cognition . i didn't get the test the free version. it ran out of free trial without sending a single prompt 😔” [source](https://twitter.com/1951686371447652352/status/2100966489813926208)
  - Complaint, Devin, r/windsurf, 2026-09-16: “since the last cli and desktop update, the service has gotten considerably worse. i have been using glm-5.2 high daily (which according to them is still free) and everything was perfect. i could use it as much as i wanted with no limits of any kind and no problems at all. but since the cli update [<strict_link> the whole service has degraded significantly. now after just a couple of prompts it rate limits me for a few minutes, which is really annoying, and after several of those pauses due to limitations i got a notice that i had used up my daily quota and can no longer use it today. what daily quota if it is supposed to be free? i do not understand. since yesterday the service has been noticeably worse, and i am thinking of canceling my subscription. there is nothing left that sets devin apart from any other service. now it is just "one more".” [source](https://www.reddit.com/r/windsurf/comments/1wi66wc/devin_cli_update_killed_the_unlimited_free_usage/)
  - Complaint, Devin, r/windsurf, 2026-09-15: “it's really embarrassing that the use of free models doesn't last long, and if you offer a free model, let people enjoy it. instead, with this move, now devin knows how many of them will cancel their subscription? then next month devin will see the drop that will happen.” [source](https://www.reddit.com/r/windsurf/comments/1whbowm/not_a_happy_camper/pa19mn2/)
  - Complaint, Devin, @cognition, 2026-09-16: “@cognition hey what happened to glm 5.2 the free tier....loved it honestly.” [source](https://twitter.com/1904957391612862464/status/2100234085218025810)

### What works

- **SWE-2 makes the Pro plan a steal** ([How much use a plan's price buys](https://feedbackbench.com/criteria/limits.plan_value.md), [Free tier and free model availability and limits](https://feedbackbench.com/criteria/billing.free_tier.md), [Single prompt, model or effort level consumes disproportionate quota](https://feedbackbench.com/criteria/limits.burn_rate.md)). Bundling unlimited SWE-2 into the entry plan turns price into Devin's loudest selling point, with users cancelling pricier rivals to switch.
  Plan value is better than peers and draws the most posts of any paying topic. Users call the entry plan the best deal in coding agents and say SWE-2 eases token anxiety. Some report cancelling top-tier plans elsewhere to keep Devin Max.
  
  One user reports a 12-hour fusion run using a small slice of weekly quota. The flip side is the window complaints above. Value is high until a limit hits.
  Evidence:
  - Praise, Devin, @cognition, 2026-09-13: “$20 for devin pro and you get free unlimited swe-2 free on @devindesktop till end of october?! this thing feels like kimi k3 on steroids, honestly i think this is the best deal in coding agents in all of 2026 and its not even close. well done @cognition team.” [source](https://twitter.com/952706191142113281/status/2099027978734727241)
  - Complaint, Devin, @cognition, 2026-09-22: “@nick_white @cognition im going to cancel cursor ultra and keep devin max next month, grok is disappointing compare to swe-2. dollar per performance is no longer there for cursor.” [source](https://twitter.com/726789031304929280/status/2102448896719147123)
  - Praise, Devin, @cognition, 2026-09-14: “to cope with token anxiety, the best choice currently is the swe-2 series model of @cognition's devin pro 🤗 <strict_link>” [source](https://twitter.com/2018156578617090049/status/2099451112717947389)
  - Praise, Devin, @cognition, 2026-09-25: “share two real cases of using devin cli: 1. running tasks with strong models + swe-2 in codex cli and devin cli, experiencing over 10 hours of uninterrupted tasks in the last two days, with failures. 2. running tasks for 12 hours using devin cli's fusion (opus-5.5 medium + swe-2 medium), only 4% of the weekly quota of the max package was used. #devin @cognition <strict_link>” [source](https://twitter.com/142110760/status/2103284980822757586)

- **Handles real engineering, not just snippets** ([Can do the user's kind of task](https://feedbackbench.com/criteria/work.capability.md), [Quality got worse or better over time](https://feedbackbench.com/criteria/models.quality_drift.md)). Users describe SWE-2 auditing whole repos, running test suites and explaining trade-offs like a careful pairing partner.
  Capability is the largest work topic and better than peers, with praise far ahead of complaints. Posts show it auditing a half-built project, running hundreds of tests and finding defects the suite missed. Others use it for recurring chores such as keeping a changelog current.
  
  Users say its communication stands out, focused on what matters. Critics say it is not frontier-grade on long, complex arcs and works best as a strong doer.
  Evidence:
  - Praise, Devin, @DevinAI, 2026-09-15: “second @devinai session may be even more impressive. i handed it a half-finished vectric aspire mcp that grok and hermes had both worked on. it audited the entire repo, ran 452 tests, found defects the suite missed, and mapped the safest path to real aspire validation. serious work.” [source](https://twitter.com/2018887236834467840/status/2099822960039112761)
  - Praise, Devin, @cognition, 2026-09-12: “amazing promo thanks @cognition day 1 of using devin*swe-2, thoughts: 🏆 swe-2 communicates what matters - refreshing 🏆 swe-2 feels like i’m pairing with a good engineer - rigour, ideas &amp; trade offs, comms 👆 🥉 swe-2 is slightly slow, but i don’t care, i’m crafting 👇 <strict_link>” [source](https://twitter.com/34404279/status/2098574999225356366)
  - Praise, Devin, @cognition, 2026-09-03: “we found it annoying to have to update our changelog, so we didn’t. but now we just have @cognition check on what we did, and update it every day. always up-to-date changelog. <strict_link>” [source](https://twitter.com/1149641989949992960/status/2095549152860176600)
  - Praise, Devin, @cognition, 2026-09-11: “it definitely isn't frontier on complex, long arc things, but it is an excellent "doer" model ime so far at the right price. i actually have my @ampcode calling devin cli to stretch my oai quota even more, but honestly, if that wasn't a major gating factor, not sure i'd be good enough to swayed, which is why i think the timing is brilliant.” [source](https://twitter.com/1699233118291361792/status/2098515309435273454)

- **Cloud sessions you can walk away from** ([Cloud and remote sandbox execution](https://feedbackbench.com/criteria/surfaces.cloud_sessions.md), [Long unattended runs and goal/loop mode](https://feedbackbench.com/criteria/work.long_running_autonomy.md)). Disposable cloud VMs with handoff to local machines let users close the laptop mid-task, and runs of 12 to 24 hours draw praise.
  Cloud sessions and long autonomy both rate better than peers, with very few complaints. Users praise teleporting a session to the cloud before a commute and picking it up later, plus SSH handoff that treats the cloud-to-local seam as a first-class feature.
  
  One user prefers Devin as a cloud harness even while keeping a rival for local work. Holdouts call blueprint setup painful and report failed SSH connections.
  Evidence:
  - Praise, Devin, @cognition, 2026-09-22: “@jkelleyrtp @cognition this lets you teleport the session quickly while you close your laptop and drive home and pick up once you’re back. great feature!” [source](https://twitter.com/618290133/status/2102188571482370518)
  - Praise, Devin, @cognition, 2026-09-22: “@cognition the ssh handoff is the wedge. most coding-agent tools died on the seam between agent vm and local machine — this treats the handoff as a first-class primitive. cloud vms are easy, the kill switch is what makes them usable.” [source](https://twitter.com/2028308162382901248/status/2102237770970546331)
  - Praise, Devin, @cognition, 2026-09-12: “@abionmorse @sherveen @devinai @cognition even without using /loop it can work for 24 hours straight it's crazy” [source](https://twitter.com/2002888897702105088/status/2098860226854220256)
  - Complaint, Devin, @cognition, 2026-09-01: “giving my first attempt to cloud agents with @cognition’s devin cloud. the blue print setup is really not easy. work tree build up is way too painful for persistent coding agents. <strict_link>” [source](https://twitter.com/218098611/status/2094576943064838315)

- **Pick any model, often on day one** ([Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md), [Automatic model routing and fallback](https://feedbackbench.com/criteria/models.routing_auto.md)). A broad, fast-updated model catalog lets users swap leads and sidekicks freely, and some cite that freedom as their reason to leave single-vendor tools.
  Model catalog access is better than peers. Users report new frontier models landing in Devin quickly, with fusion mode updated to pair a strong lead with a cheaper sidekick. One user cancelled a single-vendor subscription for the freedom to choose.
  
  Gaps remain. Users note missing computer use for SWE-2, stale models still in the list, and new installs that could not reach SWE-2 at first.
  Evidence:
  - Praise, Devin, @DevinAI, 2026-09-08: “it's official, i cancelled @claudeai for @devinai the experience so far has been much better. i cannot stand some of the limitations i run into working with claude sometimes. i really enjoy having the freedom to choose which model i want to run. devin is just nicer to use 🤷♂️ <strict_link>” [source](https://twitter.com/358932148/status/2097436804848701511)
  - Praise, Devin, @DevinAI, 2026-09-22: “fyi: gpt-6 sol, gpt-6 luna and claude opus 5.5 are already in @devinai fusion is updated too: • gpt-6 sol or opus 5.5 as your lead • gpt-6 luna as your sidekick <strict_link>” [source](https://twitter.com/415168859/status/2102491420494119103)
  - Praise, Devin, @cognition, 2026-09-22: “@cognition devin 20$ is really the best deal right now. i can swith between astra, fable, gemini 3.8 flash... (rarely use but yes you can use any model) and unlimited swe-2. moreover, codewiki and free cloud agents recently. crazy and amazing at the same time. thanks @dabit3” [source](https://twitter.com/312505543/status/2102263172736700893)
  - Complaint, Devin, @cognition, 2026-09-20: “@cognition i just downloaded your application and realized i cannot even use swe-2. you still allow your users to use 1.6(slow). if you were real confident about these benchmarks, you would allow people to try. very misleading company. you were a pioneer. now you are going to bankrupt.” [source](https://twitter.com/1437926731445321730/status/2101733059846390021)

### Under the surface

- **Orchestration is both pitch and pain** ([Subagents, parallel agents and orchestrators](https://feedbackbench.com/criteria/work.multi_agent_orchestration.md), [Automatic model routing and fallback](https://feedbackbench.com/criteria/models.routing_auto.md)). Subagents and fusion routing drive the best stories and the strangest failures, because the harness spawns or switches models without telling users.
  Users run dozens of agents in parallel and network Devins together. The same machinery shows up in complaints. One user says forced subagents dropped document quality badly until they disabled them, and even then they still spawned. Others say fan-out fell back to a default model and burned quota, or that fusion switched on silently.
  
  The common thread is invisible delegation. When it works, it feels like magic. When it misfires, users cannot see why.
  Evidence:
  - Complaint, Devin, @cognition, 2026-09-01: “@wongmjane @cognition @devindesktop watched parallel boards drift from linear within days, then agent summaries became confidently stale” [source](https://twitter.com/1803494630366785536/status/2094695949918490990)
  - Complaint, Devin, r/windsurf, 2026-09-01: “so i am using devin to create detailed documentation using skills and workflows (workflow work in cascade - not devin local). in cascade , since it does not spawn sub-agents , the documentation is of exceptional quality, adheres to all the skill based instructions. if i rate the document - 8/10 but when i am using devin local , the same document is of poor quality - if i rate it - 2/10. so in devin local, i disabled the sub-agents via overriding the config file in global (note:- it still may spawn sub-agents inspite of disabling !). the document quality improved (equal to cascade !). turns out - sub-agents does not necessarily means optimal solution! but devin forces sub-agents on us. does anyone faced similar issues ?” [source](https://www.reddit.com/r/windsurf/comments/1w4ezzy/devin_local_subagents_detoriates_the_output/)
  - Complaint, Devin, @cognition, 2026-09-14: “@januarycomputer @cognition agent in cognition that does not create a subagent inheriting the slug swe-2. i pass the model in spawn — without inherit the fan-out falls to default and the max quota disappears.” [source](https://twitter.com/326479892/status/2099470384001417470)
  - Complaint, Devin, @cognition, 2026-09-16: “@michaelkaiserai @cognition yeah i did, and used fusion accidentally. it had switched without me knowing 🥲” [source](https://twitter.com/1535774830225944576/status/2100292827221782936)

- **Promo-driven love has an expiry date** ([Free tier and free model availability and limits](https://feedbackbench.com/criteria/billing.free_tier.md), [How much use a plan's price buys](https://feedbackbench.com/criteria/limits.plan_value.md), [Quality got worse or better over time](https://feedbackbench.com/criteria/models.quality_drift.md)). Much of the current praise rests on time-limited free SWE-2, and users already ask whether pricing will hold after the promo.
  Posts that praise value often cite free or unlimited SWE-2 with a set end date. Users who lived through the free GLM tier being throttled are openly asking whether the same will happen here.
  
  Keeping free models available permanently and adding a mid-priced tier are among the top requests. How the promo ends will likely swing the paying scores more than any model release.
  Evidence:
  - Praise, Devin, @cognition, 2026-09-21: “i’m starting to test swe-2 from @cognition this week. it’s built on top of kimi k3 and tuned specifically for software engineering. for $20, you get unlimited access to the model until october 10. grok 4.7 just dropped, but honestly… i still don’t think it’s worth sleeping on swe-2.” [source](https://twitter.com/2188850740/status/2102098453107020267)
  - Praise, Devin, r/windsurf, 2026-09-11: “i started using swe2 with a pro subscription. (i can't comment on the free tier) i really liked it, especially compared to 1.7. i wrote in previous posts that i used swe 1.6 a lot, 1.7 high was a disaster, 1.7 medium was okay, but in the last 2 months i've been using glm5.2. it seemed more stable. it's important that i usually use/used devin for med. size tasks, so i'm not comparing the experience with astra or fable. if swe 2 really remains unlimited in the pro tier and there won't be enshittification, then this will be a very pleasant surprise for me. what do you think?” [source](https://www.reddit.com/r/windsurf/comments/1wd432p/swe2_first_experiences/)
  - Praise, Devin, r/CognitionLabs, 2026-09-23: “also grabbed the $20 plan to try swe-2. honestly surprised by the quality. might be even more impressive because it’s free until oct 8.” [source](https://www.reddit.com/r/CognitionLabs/comments/1wkir7w/20_devin_pro_free_swe2_my_experience_with_the/pbhu8ns/)
  - Complaint, Devin, r/windsurf, 2026-09-15: “it's really embarrassing that the use of free models doesn't last long, and if you offer a free model, let people enjoy it. instead, with this move, now devin knows how many of them will cancel their subscription? then next month devin will see the drop that will happen.” [source](https://www.reddit.com/r/windsurf/comments/1whbowm/not_a_happy_camper/pa19mn2/)

### Fine print

- Most posts come from X threads tagging Cognition or Devin accounts, which skews toward promo reactions and launch-day enthusiasm.
- The window overlaps a SWE-2 free promotion, so value and capability praise may cool once it ends.
- Many criteria, including support, onboarding and usage meter, have too few posts to rank reliably.

## Top requests

What users ask to add or change, most asked first. 480 author-weeks ask for something. Requests do not change the Feedback Score. Rule: A separate pass by Claude Sonnet 5 reads every counted post and extracts what the author asks the agent or its vendor to add or change, with the criteria it maps to and a short normalised wording; it does not touch the labels or the Feedback Score. Claude Opus 5.5 groups the wordings within each criterion (the first criterion the request maps to) into themes; code counts them. A theme counts distinct author-weeks that ask for it, per agent; across agents, one author-week per agent. Themes asked in fewer than 2 author-weeks, and requests that share no theme, are not shown. Examples: up to 3 posts per theme from different authors, without slurs, preferring posts of 60 to 450 characters, most recent first.

| Rank | Request | Criterion | Author-weeks | Posts |
|---|---|---|---|---|
| 1 | Mid-priced tier between existing plans | [How much use a plan's price buys](https://feedbackbench.com/criteria/limits.plan_value.md) | 20 | 24 |
| 2 | Official dedicated mobile app | [Mobile, remote-control and voice access](https://feedbackbench.com/criteria/surfaces.remote_mobile.md) | 16 | 16 |
| 3 | Higher overall usage limits | [How much use a plan's price buys](https://feedbackbench.com/criteria/limits.plan_value.md) | 14 | 15 |
| 4 | Keep free models available permanently | [Free tier and free model availability and limits](https://feedbackbench.com/criteria/billing.free_tier.md) | 9 | 9 |
| 5 | Free trial periods for paid plans | [Free tier and free model availability and limits](https://feedbackbench.com/criteria/billing.free_tier.md) | 8 | 9 |
| 6 | Mac VM with iOS simulator | [Computer use and browser control](https://feedbackbench.com/criteria/work.computer_browser_use.md) | 8 | 8 |
| 7 | Published exact usage limits per plan | [Pricing and plan terms stated clearly and consistently](https://feedbackbench.com/criteria/billing.pricing_clarity.md) | 7 | 8 |
| 8 | Built-in computer use capability | [Computer use and browser control](https://feedbackbench.com/criteria/work.computer_browser_use.md) | 7 | 7 |
| 9 | Bring-your-own-key support | [Connecting own API keys, local models and custom endpoints](https://feedbackbench.com/criteria/setup.provider_byok_local.md) | 6 | 8 |
| 10 | Free access to top-tier max plan | [Free tier and free model availability and limits](https://feedbackbench.com/criteria/billing.free_tier.md) | 6 | 6 |
| 11 | Preserve full context in session handoffs | [Cloud and remote sandbox execution](https://feedbackbench.com/criteria/surfaces.cloud_sessions.md) | 5 | 7 |
| 12 | Free access to specific or new models | [Free tier and free model availability and limits](https://feedbackbench.com/criteria/billing.free_tier.md) | 5 | 6 |

### 1. Mid-priced tier between existing plans

- Devin, 2026-09-23, @cognition (X): “devin from @cognition is a must, saddly for me the $20 plan are only a "taste" the limits are pretty good for the budget a single run on opus 5.5 on fusion with swe-2 gave me a 71% from daily and a 35% from my week. if in the future where i beign able to afford a max plan, well, will be a good invest! @devindesktop @cognition please make a midle tier...” [source](https://twitter.com/1747028802591494144/status/2102763773006156261)
- Devin, 2026-09-22, @cognition (X): “downgraded my cursor pro+ and added devin pro @cognition @devindesktop can we get a pro+ plan? then i'd be glad to upgrade it. swe-2 and fusion are all super great.” [source](https://twitter.com/2065820231470399488/status/2102431938682540187)
- Devin, 2026-09-22, @cognition (X): “may i ask if @cognition @devindesktop will provide a $60/month plan soon?” [source](https://twitter.com/2018156578617090049/status/2102285286172688800)

### 2. Official dedicated mobile app

- Devin, 2026-09-23, @cognition (X): “i've been an @cursor_ai user since feb 2024, but after grok 4.7, i'm looking for alternatives. @droid @cognition are in the lead for me, but what i really need is 1. cloud agents/desktop 2. agnostic harness &amp; computer use 3. mobile 4. good connectors anyone have suggestions?” [source](https://twitter.com/1902193987244408832/status/2102596587046531556)
- Devin, 2026-09-22, @cognition (X): “me waiting for a @devinai @cognition mobile app like <strict_link>” [source](https://twitter.com/1595924352/status/2102406598517772366)
- Devin, 2026-09-18, @DevinAI (X): “@bradshannon @devinai bro , need to make it accessible to mobile 📱 as well 😒 i want to play ▶️ 😫” [source](https://twitter.com/804209557706637312/status/2100932303317012557)

### 3. Higher overall usage limits

- Devin, 2026-09-25, @cognition (X): “@cognition how much usage will we get once the promotion ends 👀🧐if it’s unlimited now doesn’t that mean we could atleast have a couple of billion of tokens or 100s of dollars worth of usage” [source](https://twitter.com/860967132489801728/status/2103602614961328179)
- Devin, 2026-09-18, @DevinAI (X): “yesterday the work was intense with @devinai. there were 56 prs closed. as a result, i ran out of my devin limit and exceeded my pipeline limit on @github. we hit the bottleneck lol.” [source](https://twitter.com/326479892/status/2100934748734414957)
- Devin, 2026-09-17, @cognition (X): “is the normal quota for devin really this outrageous? running a small task with glm-5.3 directly increased the daily limit by 50% and the weekly limit? also, is the weekly limit only this much for 2 days? the computing resources are too poor. @cognition” [source](https://twitter.com/1945452790865911808/status/2100406836327547028)

### 4. Keep free models available permanently

- Devin, 2026-09-25, @cognition (X): “@cognition congrats! you deserve it! i love devin and swe-2. they’re super reliable. hope swe-2 stays free longer ;d” [source](https://twitter.com/77924805/status/2103506775735751065)
- Devin, 2026-09-24, @cognition (X): “@cognition begging to keep swe 2 free forever in cloud agents….” [source](https://twitter.com/1767985295910383616/status/2102981785785454904)
- Devin, 2026-09-22, r/windsurf (Reddit): “i also hope that after october 8th there will still be some unlimited swe model in the subscription. especially swe2 which is a very, very good deal right now.” [source](https://www.reddit.com/r/windsurf/comments/1wm5k47/swe2_is_free_until_october_8_now/pbacmzo/)

### 5. Free trial periods for paid plans

- Devin, 2026-09-24, @cognition (X): “i havent got to use devin, i think i wanna switch from grok/cursor if my experience on devin is good @devindesktop @cognition hope you guys can share me at free month devin max so i can test it....” [source](https://twitter.com/932727389523558400/status/2103072106942779813)
- Devin, 2026-09-23, @cognition (X): “@cognition @devindesktop can i get a free sub to try devin out <strict_link>” [source](https://twitter.com/2012710014277062656/status/2102631633316913401)
- Devin, 2026-09-22, @cognition (X): “@cognition if i cancel cursor pro+, can i get a free trial of devin pro？” [source](https://twitter.com/243124464/status/2102353313417351651)

### 6. Mac VM with iOS simulator

- Devin, 2026-09-17, @cognition (X): “@ptbthefirst @cognition a dedicated mac vm with ios simulator does close a real gap for mobile automation if it actually works reliably in practice.” [source](https://twitter.com/1963715144392863744/status/2100395669034852542)
- Devin, 2026-09-17, @cognition (X): “@ptbthefirst @cognition giving devin its own mac vm with an ios simulator actually solves a real bottleneck for mobile testing.” [source](https://twitter.com/2058874736470327296/status/2100391330106974685)
- Devin, 2026-09-17, @cognition (X): “@ptbthefirst @cognition devin getting a full mac vm with an ios simulator finally closes the mobile dev gap.” [source](https://twitter.com/2061512950322548736/status/2100388341631861217)

### 7. Published exact usage limits per plan

- Devin, 2026-09-17, @cognition (X): “@cognition @andrew_locke how much usage of swe2 i will have if i subscribe in max plan? it dont are clear, because theoretically it unlimited? i want more!! kkk <strict_link>” [source](https://twitter.com/1430494132368183297/status/2100588399216214137)
- Devin, 2026-09-16, @cognition (X): “@cognition , i waned to test max account, but unfortunately i have no idea of how much it would endure, which makes too hard to take a decision to risk $200 on it. how can i know about it? miss a bit of clarity on this subject. "significantly higher quotas" looks good. but the price is significantly higher too. so what does that even mean?” [source](https://twitter.com/64041638/status/2100226487110492181)
- Devin, 2026-09-11, @cognition (X): “@dabit3 @topitopongsalac @cognition and now i see why you were not replying to my message about limits. you forgot to mention the rate limits anywhere in your promos or website <strict_link>” [source](https://twitter.com/3019876798/status/2098476826360271229)

### 8. Built-in computer use capability

- Devin, 2026-09-23, @cognition (X): “i've been an @cursor_ai user since feb 2024, but after grok 4.7, i'm looking for alternatives. @droid @cognition are in the lead for me, but what i really need is 1. cloud agents/desktop 2. agnostic harness &amp; computer use 3. mobile 4. good connectors anyone have suggestions?” [source](https://twitter.com/1902193987244408832/status/2102596587046531556)
- Devin, 2026-09-23, @cognition (X): “@fei2411 @devindesktop @cognition @devinai the blogger wants the official website to properly improve its desktop version, as there is no computer use. browser automation. multiple agent sessions are still lagging.” [source](https://twitter.com/1178669733572423680/status/2102558550576992283)
- Devin, 2026-09-15, @cognition (X): “@godsboy7777 @cognition yes, exactly this, we need computer use as good as on codex and it will be unstoppable.” [source](https://twitter.com/20285011/status/2099929735056793627)

### 9. Bring-your-own-key support

- Devin, 2026-09-11, @cognition (X): “@cognition i would never pay to use a harness. insane you don’t allow byok” [source](https://twitter.com/1959797659797229568/status/2098469869536620832)
- Devin, 2026-09-10, @DevinAI (X): “hi guys @devinai can i please be able to hookup my own models via api keys. we have tokens for days.” [source](https://twitter.com/733364996722307073/status/2098124888020037687)
- Devin, 2026-09-05, @cognition (X): “@dabit3 @devinai @cognition @devindesktop please add the option to byok 🥲” [source](https://twitter.com/2078318189482737664/status/2096063333153812882)

### 10. Free access to top-tier max plan

- Devin, 2026-09-10, @cognition (X): “@joshjnunez @cognition what about a 200$ max plan to me? ahah so many giveaways but i won none :(” [source](https://twitter.com/2193149657/status/2098175761949540385)
- Devin, 2026-09-10, @cognition (X): “we’ve internally been working on building and shipping our sota image-to-image enhancement models, which make great progress already, but a free devin max plan would help us out. our results look very promising and already have the edge over tools like <strict_link>, birefnet etc. for image segmentation, but we’re pushing it until we’re highly satisfied with the models” [source](https://twitter.com/1675051807821709313/status/2098173967462756794)
- Devin, 2026-09-10, @cognition (X): “@cognition just make it free on cloud agents for max users hahahahaha” [source](https://twitter.com/17719163/status/2098070887400350089)

### 11. Preserve full context in session handoffs

- Devin, 2026-09-22, @cognition (X): “@cognition the useful feature is not remote coding by itself. it is continuity across environments: start in cli, inspect the vm over ssh, then hand the work back. the security footgun is equally clear: handoff needs explicit secret and environment boundaries, not just a nicer terminal.” [source](https://twitter.com/2051888695100514304/status/2102416772188213462)
- Devin, 2026-09-21, @cognition (X): “@cognition handoff is the right primitive. what decides whether it works is what travels with it: not just the diff, but the decisions and the dead ends from the cloud session. otherwise you sit down locally and redo work the agent already ruled out.” [source](https://twitter.com/2020874440272121856/status/2102146943623303316)
- Devin, 2026-09-21, @cognition (X): “@cognition handoffs are part of the interface. preserve commands, artifacts, tests, and rollback state or the next operator inherits a story, not evidence.” [source](https://twitter.com/2099871292480421888/status/2102126922100584665)

### 12. Free access to specific or new models

- Devin, 2026-09-18, @cognition (X): “please open free use of swe-2 only for users of devin max @cognition” [source](https://twitter.com/1011417769/status/2100826091015577762)
- Devin, 2026-09-16, @cognition (X): “@devindesktop @cognition i guess this is good, but when i use devin desktop, i immediately got used all of it lol i hope swe-2 is free in the devin desktop as well <strict_link>” [source](https://twitter.com/1828265572884467712/status/2100155249139020089)
- Devin, 2026-09-12, @cognition (X): “i'm tried @cognition on free plan. swe-1.6 slow model is the only model available, but it's not slow at all!. i wish i have access to latest model as well.” [source](https://twitter.com/1534111445985636352/status/2098665173670334911)

## Facts

| Fact | Value |
|---|---|
| Version | SWE-2 model (Medium/High/Max reasoning), Devin Fusion multi-model harness |
| Released | SWE-2: 2026 (exact date not confirmed in research) |
| Price | Free, Core/Pro $20/seat/mo, Max $200/seat/mo, Teams $80/mo base + $40/full-dev-seat, Enterprise custom. ACU (Agent Compute Unit) ~$2.25 each, ~15 min of autonomous work per ACU |
| Model | SWE-2 (enterprise API pricing $3.00/$15.00 per 1M tokens, 75% off list through 2026-12-31) |
| Surface | cloud, desktop (converging with Devin Desktop/Windsurf), IDE plugins |

## Sources

| Channel | Source | Posts |
|---|---|---|
| X | @cognition | 4111 |
| X | @DevinAI | 1984 |
| Reddit | r/windsurf | 438 |
| Reddit | Posts that name it | 113 |
| Reddit | r/CognitionLabs | 59 |
| G2 | G2 | 6 |

## Better than peers on

How much use a plan's price buys, Single prompt, model or effort level consumes disproportionate quota, Free tier and free model availability and limits, Which models are offered on a plan and when, Automatic model routing and fallback, Quality got worse or better over time, Can do the user's kind of task, Long unattended runs and goal/loop mode, Cloud and remote sandbox execution

## Worse than peers on

Short rolling usage window blocks or interrupts work, Latency, throughput and fast mode

## All 63 criteria

Criterion love: 0.5 is the category norm. n: rated author-weeks.

### Paying and limits: Better than peers (customer love 0.669, n 691)

| Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|
| [How much use a plan's price buys](https://feedbackbench.com/criteria/limits.plan_value.md) | Better than peers | 0.619 | 0.588–0.649 | 365 | 210 | 155 |
| [Free tier and free model availability and limits](https://feedbackbench.com/criteria/billing.free_tier.md) | Better than peers | 0.563 | 0.528–0.593 | 155 | 114 | 41 |
| [Single prompt, model or effort level consumes disproportionate quota](https://feedbackbench.com/criteria/limits.burn_rate.md) | Better than peers | 0.584 | 0.543–0.623 | 139 | 44 | 95 |
| [Short rolling usage window blocks or interrupts work](https://feedbackbench.com/criteria/limits.window_interrupts_work.md) | Worse than peers | 0.456 | 0.437–0.479 | 40 | 1 | 39 |
| [Pricing and plan terms stated clearly and consistently](https://feedbackbench.com/criteria/billing.pricing_clarity.md) | Typical | 0.495 | 0.457–0.540 | 38 | 3 | 35 |
| [Usage meter visibility and accuracy](https://feedbackbench.com/criteria/limits.usage_meter.md) | Too few posts | 0.499 | 0.465–0.538 | 27 | 3 | 24 |
| [Price, allowance or plan terms changed](https://feedbackbench.com/criteria/limits.allowance_change.md) | Too few posts | 0.506 | 0.473–0.542 | 20 | 2 | 18 |
| [Quota reset timing and bonus or banked resets](https://feedbackbench.com/criteria/limits.reset_schedule.md) | Too few posts | 0.503 | 0.485–0.528 | 13 | 3 | 10 |
| [Prompt cache hits, misses and invalidation](https://feedbackbench.com/criteria/limits.prompt_cache.md) | Too few posts | 0.506 | 0.490–0.520 | 13 | 9 | 4 |
| [Using an existing subscription across tools](https://feedbackbench.com/criteria/billing.subscription_portability.md) | Too few posts | 0.493 | 0.478–0.509 | 11 | 2 | 9 |
| [Pay-as-you-go overage, fallback billing and spend caps](https://feedbackbench.com/criteria/billing.overage_charges.md) | Too few posts | 0.495 | 0.489–0.499 | 4 | 0 | 4 |

Most recent posts:

- Praise, 2026-09-27, r/windsurf (Reddit): “they had abandoned since almost a year ago. it was up until a month ago, but it didn't even work when it was up. you could install devin desktop and enjoy your unlimited free tab complete there.” [source](https://www.reddit.com/r/windsurf/comments/1wr7p7l/why_isnt_windsurf_available_in_vscode_anymore/pceq77b/)
- Praise, 2026-09-27, @cognition (X): “been 96 hours. @cognition @devindesktop i've never seen a harness as good as this. they don't even have the /goal, but their execution surpass any harness with /goal on the market. it's f*cking mind blowing. 20$, unlimited swe 2 you should definitely give it a shot <strict_link> <strict_link>” [source](https://twitter.com/1825243355501973504/status/2104016524725952607)
- Praise, 2026-09-27, @cognition (X): “@max18martin @cognition i have to say that swe-2 can even be used in many scenarios to match gpt-6 sol, and the former is surprisingly free for pro plans and above.” [source](https://twitter.com/1859530427905736704/status/2104025642329330144)
- Complaint, 2026-09-27, @cognition (X): “devin from @cognition is surely a great workhorse but it burnt through my entire weekly quota in less than 8 hours 🤯” [source](https://twitter.com/15290915/status/2104239118419112032)
- Complaint, 2026-09-27, @cognition (X): “@ditlied @cognition @devindesktop been polite is something important in life. anyway, it's cheaper, not free” [source](https://twitter.com/1825243355501973504/status/2104286835711062127)
- Complaint, 2026-09-27, @cognition (X): “so daily * 2 = weekly @devindesktop @cognition ? whats the point of calling it weekly quota? might as well call it 2-day quota. not good. <strict_link>” [source](https://twitter.com/1299192048843517953/status/2104289857044640200)

### Setting up and connecting: Typical (customer love 0.511, n 92)

| Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|
| [Install, launch and sign-in](https://feedbackbench.com/criteria/setup.install_signin.md) | Too few posts | 0.556 | 0.522–0.589 | 25 | 13 | 12 |
| [MCP servers, plugins, skills and hooks](https://feedbackbench.com/criteria/setup.extensions_mcp.md) | Too few posts | 0.502 | 0.481–0.523 | 25 | 13 | 12 |
| [Onboarding, discoverability and documentation](https://feedbackbench.com/criteria/setup.onboarding_docs.md) | Too few posts | 0.499 | 0.477–0.524 | 17 | 3 | 14 |
| [IDE and editor integration](https://feedbackbench.com/criteria/setup.ide_integration.md) | Too few posts | 0.515 | 0.497–0.534 | 16 | 10 | 6 |
| [Connecting own API keys, local models and custom endpoints](https://feedbackbench.com/criteria/setup.provider_byok_local.md) | Too few posts | 0.473 | 0.455–0.489 | 13 | 1 | 12 |

Most recent posts:

- Praise, 2026-09-27, r/windsurf (Reddit): “gemini is dumb codeium is like 2 years old..windsurf went away about 3 months ago and is now devin desktop. they stopped supporting the vs code extension about a month ago. but devin desktop is a vs code fork. just download the ide, it's going to feel very native for you” [source](https://www.reddit.com/r/windsurf/comments/1wr7p7l/why_isnt_windsurf_available_in_vscode_anymore/pcap3zm/)
- Praise, 2026-09-27, r/windsurf (Reddit): “devin has invested a lot in controlling how the ais think inside the ide, to maximize context and code management tools. if you are going to pay for devin, try the devin desktop ide. it's a fork of vs code so it's not unfamiliar. the ide is good for almost any workflow, from lightning-fast code completion to highly reliable vibe coding with swe-2.” [source](https://www.reddit.com/r/windsurf/comments/1wr7p7l/why_isnt_windsurf_available_in_vscode_anymore/pcb72fn/)
- Praise, 2026-09-25, @cognition (X): “@matthewcp tbh @cognition’s devinwiki is really good for this” [source](https://twitter.com/1595924352/status/2103536577209094392)
- Complaint, 2026-09-27, @DevinAI (X): “@devinai guys no offence but why when i open the website i cant tell what the product does right away, this is just a feedback that you guys could be getting more users if the main example is suffience to explain it, best” [source](https://twitter.com/1657919466498412546/status/2104085645853331896)
- Complaint, 2026-09-25, @cognition (X): “can someone from @cognition tell me why i have to use the browser instead of desktop app for wiki sessions and session ios simulator 🫩” [source](https://twitter.com/4108256724/status/2103280180357947496)
- Complaint, 2026-09-25, @cognition (X): “@cognition is there a way to get devin to respond to our agent on slack? trying to get them to talk to eachother. devin’s suggestion here didn’t work. <strict_link>” [source](https://twitter.com/787864/status/2103318326864957759)

### Choosing models: Better than peers (customer love 0.636, n 176)

| Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|
| [Quality got worse or better over time](https://feedbackbench.com/criteria/models.quality_drift.md) | Better than peers | 0.551 | 0.520–0.583 | 66 | 32 | 34 |
| [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | Better than peers | 0.600 | 0.567–0.633 | 62 | 43 | 19 |
| [Automatic model routing and fallback](https://feedbackbench.com/criteria/models.routing_auto.md) | Better than peers | 0.559 | 0.529–0.588 | 39 | 25 | 14 |
| [Reasoning effort setting and its defaults](https://feedbackbench.com/criteria/models.effort_control.md) | Too few posts | 0.507 | 0.489–0.525 | 17 | 9 | 8 |

Most recent posts:

- Praise, 2026-09-27, @cognition (X): “wtf is that pareto??? did swe 2 just fucking break it??? @cognition @devindesktop i can't wait for swe 3!!! your team is crazy!!! <strict_link> <strict_link>” [source](https://twitter.com/1825243355501973504/status/2104017046799372411)
- Praise, 2026-09-27, @cognition (X): “@bnistordev @learnmore_smart @cognition @devindesktop opus 5.5 + swe-2 cheaper and even better” [source](https://twitter.com/17719163/status/2104128141895643174)
- Praise, 2026-09-27, @cognition (X): “@gamerz_artist try @cognition ... good usage for all models” [source](https://twitter.com/2080910665041010688/status/2104299903010672820)
- Complaint, 2026-09-27, @DevinAI (X): “@markfenner @devinai i suspect they have used a quantized version causing the models iq to drop” [source](https://twitter.com/1916897001922506752/status/2104092718213583286)
- Complaint, 2026-09-25, @DevinAI (X): “@notjazii @devinai yeah fr they shouldn't nerf it” [source](https://twitter.com/2012475539324559360/status/2103554876320137216)
- Complaint, 2026-09-25, @DevinAI (X): “@notjazii @devinai don't worry, they won't nerf it down until the release of opus 5.6” [source](https://twitter.com/1388715421864402947/status/2103563089996337471)

### Instructing and context: Typical (customer love 0.500, n 46)

| Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|
| [Finding the right files in the codebase](https://feedbackbench.com/criteria/context.codebase_retrieval.md) | Too few posts | 0.505 | 0.490–0.522 | 13 | 8 | 5 |
| [Memory and state carried across sessions](https://feedbackbench.com/criteria/context.session_memory.md) | Too few posts | 0.498 | 0.483–0.512 | 10 | 5 | 5 |
| [Direct in-prompt instructions and caps are followed](https://feedbackbench.com/criteria/context.instruction_following.md) | Too few posts | 0.501 | 0.488–0.517 | 9 | 3 | 6 |
| [Asks the user versus guessing](https://feedbackbench.com/criteria/context.clarifying_questions.md) | Too few posts | 0.492 | 0.485–0.498 | 4 | 0 | 4 |
| [Output degrades as the context window fills](https://feedbackbench.com/criteria/context.long_context_decay.md) | Too few posts | 0.494 | 0.488–0.499 | 4 | 0 | 4 |
| [Context compaction keeps what matters, cheaply and quickly](https://feedbackbench.com/criteria/context.compaction.md) | Too few posts | 0.499 | 0.492–0.509 | 3 | 1 | 2 |
| [Images, PDFs and file attachments as input](https://feedbackbench.com/criteria/context.attachments.md) | Too few posts | 0.506 | 0.500–0.516 | 2 | 2 | 0 |
| [Persistent project rules files are read and obeyed](https://feedbackbench.com/criteria/context.instruction_files.md) | Too few posts | 0.497 | 0.492–0.500 | 1 | 0 | 1 |

Most recent posts:

- Praise, 2026-09-26, @cognition (X): “@marvinvonhagen @cognition seeing 200m messages exchanged reminds me of when i switched to a tool that remembered everything and it changed how i work.” [source](https://twitter.com/332239817/status/2103772838788300834)
- Praise, 2026-09-23, @cognition (X): “@cognition direct messages, file attachments and self-updating make this feel like more than a code window — it can handle the handoffs around the task too” [source](https://twitter.com/1037725470630891520/status/2102827914614202471)
- Praise, 2026-09-21, r/codex (Reddit): “you probably got a quantized model. tell it to provide a handoff and start over. but before you do that, get a second opinion from swe-2 or deepseek. you'd also probably get better results if you just used swe-2 and told it to call codex cli astra as an advisor. devin doesn't block itself on tests and such so much and does what you ask.” [source](https://www.reddit.com/r/codex/comments/1wm88ng/stuck_in_the_mud_spinning_the_wheels_but_no/pb4tx7q/)
- Complaint, 2026-09-25, @cognition (X): “@cognition devin just hit the billion dollar mark and still probably asks for a clearer ticket 😂” [source](https://twitter.com/1699417980155637761/status/2103506346553610593)
- Complaint, 2026-09-22, @cognition (X): “@cognition you're losing all the context right? different environments” [source](https://twitter.com/1862977676136337408/status/2102319987910451638)
- Complaint, 2026-09-19, @cognition (X): “@brandon_galang @cognition @devinai one thing i really like about it is it tells you the size of your thread and how many acu its currently cost so u can change to a new thread. it seems like they dont have compaction on cloud” [source](https://twitter.com/1665450872363708417/status/2101143353551388871)

### Doing the work: Better than peers (customer love 0.602, n 620)

| Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|
| [Can do the user's kind of task](https://feedbackbench.com/criteria/work.capability.md) | Better than peers | 0.554 | 0.522–0.587 | 449 | 330 | 119 |
| [Subagents, parallel agents and orchestrators](https://feedbackbench.com/criteria/work.multi_agent_orchestration.md) | Typical | 0.487 | 0.456–0.519 | 82 | 47 | 35 |
| [Long unattended runs and goal/loop mode](https://feedbackbench.com/criteria/work.long_running_autonomy.md) | Better than peers | 0.539 | 0.511–0.563 | 47 | 43 | 4 |
| [Computer use and browser control](https://feedbackbench.com/criteria/work.computer_browser_use.md) | Too few posts | 0.501 | 0.480–0.522 | 27 | 16 | 11 |
| [Does unrequested work or over-engineers](https://feedbackbench.com/criteria/work.scope_overreach.md) | Too few posts | 0.524 | 0.487–0.558 | 13 | 3 | 10 |
| [Spins, loops or gets stuck without progress](https://feedbackbench.com/criteria/work.stuck_loops.md) | Too few posts | 0.517 | 0.485–0.563 | 12 | 2 | 10 |
| [Length and clarity of replies, summaries and comments](https://feedbackbench.com/criteria/work.response_verbosity.md) | Too few posts | 0.509 | 0.493–0.526 | 12 | 6 | 6 |
| [Breaks existing code or reintroduces bugs](https://feedbackbench.com/criteria/work.regressions_introduced.md) | Too few posts | 0.490 | 0.478–0.503 | 10 | 1 | 9 |
| [Tool approval prompts and autonomy modes](https://feedbackbench.com/criteria/work.permission_prompts.md) | Too few posts | 0.492 | 0.481–0.508 | 9 | 1 | 8 |
| [Frontend and visual UI output](https://feedbackbench.com/criteria/work.frontend_ui.md) | Too few posts | 0.493 | 0.479–0.507 | 8 | 4 | 4 |
| [Risky or irreversible actions without confirmation](https://feedbackbench.com/criteria/work.destructive_actions.md) | Too few posts | 0.499 | 0.485–0.517 | 7 | 1 | 6 |
| [Stops mid-task or answers instead of acting](https://feedbackbench.com/criteria/work.premature_stop.md) | Too few posts | 0.490 | 0.482–0.498 | 5 | 0 | 5 |
| [Diagnosing and fixing reported bugs](https://feedbackbench.com/criteria/work.bug_diagnosis.md) | Too few posts | 0.507 | 0.502–0.513 | 4 | 4 | 0 |
| [Git commits, branches and sync](https://feedbackbench.com/criteria/work.git_workflow.md) | Too few posts | 0.502 | 0.495–0.511 | 2 | 1 | 1 |
| [Plan-before-edit mode](https://feedbackbench.com/criteria/work.plan_mode.md) | Too few posts | 0.505 | 0.500–0.515 | 2 | 2 | 0 |
| [Games checks instead of fixing the problem](https://feedbackbench.com/criteria/work.reward_hacking.md) | Too few posts | 0.499 | 0.496–0.500 | 1 | 0 | 1 |
| [Safety filters block legitimate coding tasks](https://feedbackbench.com/criteria/work.safety_refusals.md) | Too few posts | 0.499 | 0.496–0.500 | 1 | 0 | 1 |
| [Caves to or argues with the user's judgement](https://feedbackbench.com/criteria/work.sycophancy_pushback.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 |

Most recent posts:

- Praise, 2026-09-27, r/windsurf (Reddit): “devin has invested a lot in controlling how the ais think inside the ide, to maximize context and code management tools. if you are going to pay for devin, try the devin desktop ide. it's a fork of vs code so it's not unfamiliar. the ide is good for almost any workflow, from lightning-fast code completion to highly reliable vibe coding with swe-2.” [source](https://www.reddit.com/r/windsurf/comments/1wr7p7l/why_isnt_windsurf_available_in_vscode_anymore/pcb72fn/)
- Praise, 2026-09-27, r/windsurf (Reddit): “they are not in the same class of capability. copilot is not a very good harness” [source](https://www.reddit.com/r/windsurf/comments/1wr7p7l/why_isnt_windsurf_available_in_vscode_anymore/pcd4mf5/)
- Praise, 2026-09-27, @cognition (X): “if you have claude subscription and @cognition cloud agents. try it it is amazing combo, it ships like crazy when you sleep. @cognition swe 2 max is a really good model and on cloud, it has macos, can code, run test, take screenshot and record evidence. opus 5.5 is really well as orchestrator, review and merge pr on your machine. you can feed opus (in claude code or hermes) devin api key, it knows…” [source](https://twitter.com/1767985295910383616/status/2104014647724761554)
- Complaint, 2026-09-27, @cognition (X): “another task going to be 24 hours! @devinai @devindesktop @cognition you guys are cooking my projects!!! unlimited swe 2 with a 20usd plan, are you kidding me? that is literally the best coding plan in the world <strict_link>” [source](https://twitter.com/1825243355501973504/status/2104262892090781744)
- Complaint, 2026-09-27, @DevinAI (X): “@doodlestein @hraness @chatgpt @devinai @claudeai @cursor_ai @bot @zeddotdev the part i’d watch is observability. more terminals only helps if you can see which run stalled, what changed, and whether the result is safe to merge. otherwise it’s just a very expensive wall of tabs.” [source](https://twitter.com/2103910196330242049/status/2104140047633326358)
- Complaint, 2026-09-27, @DevinAI (X): “@doodlestein @hraness @chatgpt @devinai @claudeai @cursor_ai @bot @zeddotdev sixty-four accounts. at that point the agents are managing you, not the other way around.” [source](https://twitter.com/2087402808756629504/status/2104147885679845800)

### Checking and finishing: Typical (customer love 0.494, n 61)

| Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|
| [Builds, tests or runs its own changes](https://feedbackbench.com/criteria/verify.self_testing.md) | Too few posts | 0.514 | 0.495–0.532 | 23 | 18 | 5 |
| [Reviewing and approving the agent's changes](https://feedbackbench.com/criteria/verify.change_review_ui.md) | Too few posts | 0.473 | 0.456–0.490 | 21 | 3 | 18 |
| [Agent-performed code review finds real issues](https://feedbackbench.com/criteria/verify.agent_code_review.md) | Too few posts | 0.500 | 0.481–0.515 | 10 | 8 | 2 |
| [Claims work is done or fixed when it is not](https://feedbackbench.com/criteria/verify.false_completion.md) | Too few posts | 0.489 | 0.482–0.496 | 8 | 0 | 8 |

Most recent posts:

- Praise, 2026-09-27, @DevinAI (X): “@devinai just cooked. it tested itself and shipped me an actual video i could watch. momentum v0.3.0 is a full frontend + architecture reset. gtm target: end of october. get in. @tacticocc <strict_link>” [source](https://twitter.com/1263379788246347776/status/2104233763060527444)
- Praise, 2026-09-25, @cognition (X): “@cognition i was a hater, but i'm now using deepwiki and devin reviews on ci a lot. so congrats.” [source](https://twitter.com/1458111452397674503/status/2103515665168470122)
- Praise, 2026-09-23, @cognition (X): “ok @cognition devin is the code reviewer you want checking everything. and great value.” [source](https://twitter.com/1440778796727091206/status/2102831105569378664)
- Complaint, 2026-09-27, @DevinAI (X): “@ryancarson @devinai @linear @hellountangle zero local dev just relocates the humans to the one place that still matters: review. which makes the reviewer the production line — and the only one holding the loss when the diff reads fine and isn't.” [source](https://twitter.com/2065683882587144192/status/2104001307874885845)
- Complaint, 2026-09-27, @DevinAI (X): “@hraness @chatgpt @devinai @claudeai @cursor_ai @bot @zeddotdev 39 terminals make human review the bottleneck, not code generation” [source](https://twitter.com/1803494630366785536/status/2104144285939404963)
- Complaint, 2026-09-26, @cognition (X): “@cognition @devindesktop can you please add a capability for an agent to set up a schedule? or let your agent know that it doesn't have that capability instead of a false promise <strict_link>” [source](https://twitter.com/383156096/status/2103726198362873948)

### Interface and sessions: Better than peers (customer love 0.592, n 157)

| Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|
| [Cloud and remote sandbox execution](https://feedbackbench.com/criteria/surfaces.cloud_sessions.md) | Better than peers | 0.571 | 0.541–0.599 | 88 | 77 | 11 |
| [How the interface shows work, and what the user can configure](https://feedbackbench.com/criteria/ui.display_settings.md) | Typical | 0.508 | 0.482–0.537 | 47 | 19 | 28 |
| [Mobile, remote-control and voice access](https://feedbackbench.com/criteria/surfaces.remote_mobile.md) | Too few posts | 0.495 | 0.477–0.515 | 21 | 9 | 12 |
| [Saving, switching, resuming and rewinding sessions](https://feedbackbench.com/criteria/ui.session_history.md) | Too few posts | 0.487 | 0.475–0.501 | 11 | 1 | 10 |
| [Stopping and steering a running agent](https://feedbackbench.com/criteria/ui.interrupt_steer.md) | Too few posts | 0.500 | 0.489–0.511 | 6 | 3 | 3 |

Most recent posts:

- Praise, 2026-09-26, @cognition (X): “this is 100% available in devin cloud sessions, it works very well, and we've had it for quite some time, we don't support it locally as we can't ensure the user's machine will be available or on. you can ask in natural language or choose from many different templates. the local agent should understand this so we'll look into it.” [source](https://twitter.com/17189394/status/2103877279424094668)
- Praise, 2026-09-25, @cognition (X): “one of my favorite features of @cognition devin is their cloud agents and how they can test changes for you this was a session i had while on my way to university this morning. incredibly useful! <strict_link>” [source](https://twitter.com/3293793720/status/2103276385481654569)
- Praise, 2026-09-25, @DevinAI (X): “@adelwu_ @devinai i'm saying the design change was sexy” [source](https://twitter.com/1905803078135336960/status/2103284719832203686)
- Complaint, 2026-09-27, @cognition (X): “@cognition @devindesktop can you give more control on the session panel. for example between windows itd be great to hide or minimize sessions not relevant for that window in order to concentrate on separate domains seamlessly por favor” [source](https://twitter.com/1176601698481201152/status/2104032779713315233)
- Complaint, 2026-09-26, @cognition (X): “@cognition when you click new session in a new 'tab' can you fix making the new session opening in place vs a separately active tab which requires moving the session as a next step por favor” [source](https://twitter.com/1176601698481201152/status/2103761772893384738)
- Complaint, 2026-09-26, @cognition (X): “@shayanshafii @cognition dang nothing for remote?” [source](https://twitter.com/1974384377770504194/status/2103919366337396816)

### Reliability and speed: Typical (customer love 0.533, n 149)

| Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|
| [Latency, throughput and fast mode](https://feedbackbench.com/criteria/rel.response_speed.md) | Worse than peers | 0.468 | 0.438–0.497 | 70 | 21 | 49 |
| [Client crashes, freezes and failed tool execution](https://feedbackbench.com/criteria/rel.client_failures.md) | Typical | 0.497 | 0.446–0.557 | 50 | 3 | 47 |
| [Outages, server errors and capacity or rate errors](https://feedbackbench.com/criteria/rel.service_errors.md) | Too few posts | 0.589 | 0.527–0.644 | 27 | 8 | 19 |
| [Updates break working setups](https://feedbackbench.com/criteria/rel.update_breakage.md) | Too few posts | 0.533 | 0.501–0.567 | 14 | 6 | 8 |

Most recent posts:

- Praise, 2026-09-27, @DevinAI (X): “@hraness @chatgpt @devinai me too and i'm finally getting satisfied after 6 mos of tinkering. especially around reliability and mutli host orchestration. in fact there's so much to orchestrate not just agents.” [source](https://twitter.com/1887172409125658624/status/2104342859843273166)
- Praise, 2026-09-26, @cognition (X): “@agentmasterkey @devinai @cognition @openai @thsottiaux best harness is the one that keeps working when codex is down. failover is the feature” [source](https://twitter.com/2009223361969442816/status/2103644784011407861)
- Praise, 2026-09-23, @DevinAI (X): “@guybedo @devinai @zeddotdev it's been pretty snappy for me” [source](https://twitter.com/896906084014845952/status/2102564705009275354)
- Complaint, 2026-09-27, r/CognitionLabs (Reddit): “looks like there is a bug with the devin tab. im using a m4 macbook air 24gb. the problem started with golden gate update. i asked claude opus 5.5 and after debugging it found this: you're right, it isn't normal. i found the code path responsible, and it's a performance bug inside devin's built-in extension, not something in your setup. **where the 2.2s goes** (newest profile, `exthost-b13005.cpup…” [source](https://www.reddit.com/r/CognitionLabs/comments/1w23p1r/devin_is_simply_too_slow/pcfa29z/)
- Complaint, 2026-09-27, @DevinAI (X): “@markfenner @devinai why devin take so much time while building?” [source](https://twitter.com/2278324309/status/2104309916702048679)
- Complaint, 2026-09-26, @cognition (X): “@dabit3 i must say swe-2 is still slow but its also really good, thank you for this @cognition” [source](https://twitter.com/1973083865607708673/status/2103901813158355257)

### Account and support: Too few posts (customer love 0.522, n 29)

| Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|
| [Support, refunds and issue handling](https://feedbackbench.com/criteria/account.support.md) | Too few posts | 0.512 | 0.485–0.537 | 19 | 5 | 14 |
| [Wrong charges, failed payments and plan provisioning](https://feedbackbench.com/criteria/account.billing_errors.md) | Too few posts | 0.489 | 0.482–0.495 | 9 | 0 | 9 |
| [Data retention, training use and deployment isolation](https://feedbackbench.com/criteria/account.data_privacy.md) | Too few posts | 0.508 | 0.496–0.526 | 3 | 2 | 1 |
| [Account bans and access restrictions](https://feedbackbench.com/criteria/account.bans_restrictions.md) | Too few posts | 0.499 | 0.495–0.500 | 1 | 0 | 1 |

Most recent posts:

- Praise, 2026-09-27, r/CognitionLabs (Reddit): “cloud is not accesible to others” [source](https://www.reddit.com/r/CognitionLabs/comments/1wq4noh/where_to_install_devin_desktop/pcfcehn/)
- Praise, 2026-09-26, @cognition (X): “@dabit3 @cognition @devindesktop ah, got it. thanks for the prompt response” [source](https://twitter.com/383156096/status/2103904109019762779)
- Praise, 2026-09-25, @cognition (X): “@cognition congrats! first class customer success service and ultra high agency! 🧡” [source](https://twitter.com/1079053150634602496/status/2103507609454010767)
- Complaint, 2026-09-24, r/windsurf (Reddit): “my bank notified me of unusual charges from windsurf/devin. i see all of my on-demand balance used up and multiple charges for on-demand use. i was an early user of windsurf and had extra use credit built up before they started used limits. i havn't used devin for 1.5 weeks. i check my code and repositories to find no changes since 1.5 weeks ago. checked devin website to find no history of prompts…” [source](https://www.reddit.com/r/windsurf/comments/1wpeoew/unauthorized_and_unaccounted_token_usage_and/)
- Complaint, 2026-09-23, @cognition (X): “@devindesktop @cognition there support sucks and the payment is not even working and support form not workign either tried on 3 devices 3 accounts 3 working cards and still payment not working. <strict_link>” [source](https://twitter.com/1900952240342659072/status/2102575178966356397)
- Complaint, 2026-09-23, @cognition (X): “@chris_wozniczek @dabit3 @cognition hey i have a error buying a sub from devin but it keeps failing on 3 devices and 3 working cards on 3 different accounts support said they cant do anything about it! please help me here! i also asked for a free on in the image and they have yet to reply its been over 5 days! <strict_link>” [source](https://twitter.com/1900952240342659072/status/2102584861236363750)
