# Pi (Earendil Works (open source))

Feedback Bench, coding agents, built 2026-10-01, window 2026-08-31 to 2026-09-27. Web page: https://feedbackbench.com/#/agent/pi

| Measure | Value |
|---|---|
| Rank | 7 of 17 (rank range 7–7) |
| Feedback Score | 52.7 (95% interval 51.9–53.3) |
| Popularity | 0.456 (share of voice 2.77%) |
| Customer love | 0.608 (95% interval 0.591–0.623) |
| Top quadrant | no |
| Authors | 2732 |
| Posts counted | 6213 |
| Posts that judge the agent | 2146 |
| Criteria better / worse than peers | 14 / 0 of 63 |

## The brief

Written by Claude Opus 5.5 from 110 labelled posts and the numbers on this page. Interpretation, not measurement: every quote is verbatim and links to its post.

**A lean harness you own, if you build the rest yourself.**

TL;DR:

- Users pick Pi for its extension model and tiny context footprint, then reshape it per project.
- Subagents, stronger compaction and session persistence mostly arrive through community extensions, not the core.
- The loudest request is bringing an existing subscription, especially Claude, into Pi.

### What hurts

- **Subagents are a do-it-yourself project** ([Subagents, parallel agents and orchestrators](https://feedbackbench.com/criteria/work.multi_agent_orchestration.md)). Pi ships without native subagents, so users wire orchestration from extensions and external tools, and coordination stays loose.
  Posts describe subagents arriving only through extensions, and one user built a separate manager after finding agents 'too loosely connected' with weak lifecycle control. Others juggle concurrent work in the terminal.
  
  A built-in orchestrator mode and a live subagent dashboard both show up among user requests. Some users doubt subagents help at all and want targeted jobs instead, such as a simplify pass on a different model.
  Evidence:
  - Complaint, Pi, @pidotdev, 2026-09-02: “@pidotdev doesn't have subagents, but i use an extension. this, with gpt-5.6, spawns subagents like there is crazy experiment: can pi fix itself? asked pi to create an extension that appends the system prompt on gpt-5.6 + used @mattpocockuk 's skill to write a succinct prompt <strict_link>” [source](https://twitter.com/1153490267938275328/status/2095050463020552321)
  - Complaint, Pi, r/PiCodingAgent, 2026-09-23: “i actually used pi-subagents + pi-intercom before building [pi herdsman](<strict_link>). they got me pretty close, but i kept running into the same kind of thing mentioned here: agents felt a bit too loosely connected and coordination/lifecycle wasn't as tight as i wanted. herdsman is my take on that setup: herdr owns the actual pi sessions, while herdsman manages delegation, ownership, questions/results, and cleanup around them. the lead stays interactive and keeps the bigger-picture context, while the agents do the detailed work separately. it does require herdr tough. [<strict_link>” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wogo23/has_anyone_tried_piintercom/pbn4ezj/)
  - Complaint, Pi, r/PiCodingAgent, 2026-09-12: “i was on terminal for quite a while, for me the two major issues were managing lots of concurrent work, and terminal limitations around correctly displaying bidi text” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wdjlrf/farcaster_pi_neovim_in_a_desktop_app/p9b56r4/)
  - Complaint, Pi, r/PiCodingAgent, 2026-09-14: “i'm not convinced that subagents have much value, and quite possibly they cause harm. <strict_link> what i do want better tools for is specific jobs. eg. always running a simplify loop at the end of task, but doing it with a different model or different thinking level.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wfvpzk/i_stopped_treating_subagents_as_disposable/p9q2jwa/)

- **Compaction stalls and loses the thread** ([Context compaction keeps what matters, cheaply and quickly](https://feedbackbench.com/criteria/context.compaction.md), [Output degrades as the context window fills](https://feedbackbench.com/criteria/context.long_context_decay.md)). Native compaction halts runs or degrades output as the window fills, so users cap compactions or swap in third-party versions.
  One user says Pi stops after compaction and needs a manual nudge, and never lets it compact more than twice. Another reports compaction not working at all, with performance sliding near the context limit. A fork exists because the original blocked model switching after native compaction.
  
  Better compaction summaries are a recurring request. The fixes users praise are extensions, not the default.
  Evidence:
  - Complaint, Pi, r/PiCodingAgent, 2026-09-11: “i'm using q4_0 kv with 181 ctx, using 3.8 27b q3. it just moves right along. it will stop after compaction and i just tell it to continue. i never let it compact more than twice.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wdi95m/session_eventually_get_stuck_when_using_a_small/p96xuek/)
  - Complaint, Pi, r/opencodeCLI, 2026-09-03: “in my case muse is really awesome. i plan all the jobs with a more intelligent model (opus, sol or even terra in some cases) then i execute the plan with muse. the problem with muse i noticed is when you are not very specific with the spec and it starts to do things because you weren’t intend to do, also for me in pi agent compactions is not working for some reason (then you get closer to the context window, it stars to be less performant).” [source](https://www.reddit.com/r/opencodeCLI/comments/1vu8mi0/better_in_your_experience_mimo_vs_hy3_vs_muse/p7jp6lj/)
  - Complaint, Pi, r/PiCodingAgent, 2026-09-27: “<strict_link> or you can use the original, i've made a fork so it works with switching to other models, original doesn't allow you to switch the model after native compaction” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wrq50c/new_session_handoff_vs_compact_which_do_you_prefer/pcewx2q/)
  - Complaint, Pi, @pidotdev, 2026-09-22: “@pidotdev remote control native support. better compaction. human-writing like :)” [source](https://twitter.com/89646598/status/2102407468450296138)

- **The TUI freezes and yanks the view** ([Client crashes, freezes and failed tool execution](https://feedbackbench.com/criteria/rel.client_failures.md), [How the interface shows work, and what the user can configure](https://feedbackbench.com/criteria/ui.display_settings.md)). The interface pauses after tool calls and force-scrolls to the latest call, interrupting anyone reading back through a transcript.
  Users report a roughly one-second UI freeze each time a tool call finishes. Tool calls also pull the view to the present while you scroll older output, and fullscreen mode is described as too slow to scroll.
  
  Other posts mention piles of lingering node processes and relative file links that will not open, which one user patched with a display-time extension.
  Evidence:
  - Complaint, Pi, @pidotdev, 2026-09-22: “@pidotdev could you improve the performance? every time the agent finishes a tool call, the ui freezes for about a second and i can't do anything” [source](https://twitter.com/1601235888029446144/status/2102387996079317156)
  - Complaint, Pi, @pidotdev, 2026-09-25: “hey @pidotdev, any reason why tool calls force-scroll the view? if i'm reading something older in the transcript, tool calls will pull me right to the present tool call. i use regular instead of fullscreen because fullscreen is just too slow to scroll with the mouse” [source](https://twitter.com/1661172519989116928/status/2103546890469986704)
  - Complaint, Pi, r/PiCodingAgent, 2026-09-26: “every time i turn on my mac and see so many node processes, it’s really terrible. i think i‘ll try.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wqe2v1/meet_pig_the_pi_coding_harness_that_is_yours_but/pc4svol/)
  - Complaint, Pi, @pidotdev, 2026-09-20: “relative file link may not open in pi cli or tui. that's why i made a really simple pi-local-file-links to resolve links at display time only. session history and model context stay unchanged. <strict_link> <strict_link> @pidotdev #pi #codingagent” [source](https://twitter.com/1757494482579451904/status/2101584398336561630)

- **Claude subscriptions stay locked out** ([Using an existing subscription across tools](https://feedbackbench.com/criteria/billing.subscription_portability.md), [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md)). Users cannot run their Claude plan inside Pi, so they route to OpenAI subscriptions or pay per token for Anthropic models.
  This is the single most requested change. Users who want Opus in Pi say the lockout makes them angry, and some settle for alternatives they call mediocre by comparison.
  
  The workaround pattern is consistent. Run OpenAI subscriptions in Pi, keep Anthropic usage light and metered, and lobby the model vendor directly.
  Evidence:
  - Complaint, Pi, @pidotdev, 2026-09-04: “@trq212 @pidotdev @badlogicgames joining the chorus — then let us use sub in pi” [source](https://twitter.com/140337831/status/2095784426315960517)
  - Praise, Pi, @pidotdev, 2026-09-14: “@avindrafernando @pidotdev the gpt models have just been working better for me. the fact that i can't use my claude max subscription with pi also makes me very angry!” [source](https://twitter.com/11069822/status/2099537485168599225)
  - Complaint, Pi, r/PiCodingAgent, 2026-09-25: “same. it's an inescapable fact that anthropic doesn't want their subscriptions used with other agents. so i use openai's subscriptions and make light use of pay-per-token anthropic apis. someone downvoted you, but this was the only logical way for me.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wq1bv2/whats_the_best_option_for_using_anthropic/pc0oobo/)
  - Complaint, Pi, @pidotdev, 2026-09-23: “@manuel_kehl @iannuttall @pidotdev precisely but they aren’t as good as opus 5.5 so yeah you use mediocre” [source](https://twitter.com/1529883822044717057/status/2102690908697509992)

- **Minimalism shifts the work onto you** ([Can do the user's kind of task](https://feedbackbench.com/criteria/work.capability.md), [How the interface shows work, and what the user can configure](https://feedbackbench.com/criteria/ui.display_settings.md), [Onboarding, discoverability and documentation](https://feedbackbench.com/criteria/setup.onboarding_docs.md)). The bare core forces setup and tuning time, and some users say heavier harnesses beat it out of the box.
  A joking post captures it: all tokens go into customizing Pi before any real work happens. Critics point to benchmarks where tool-heavy harnesses outscore Pi, and one user finds a specific model lazier in Pi than in its native harness.
  
  The minimalism that fans love is the same thing that leaves gaps for newcomers.
  Evidence:
  - Complaint, Pi, r/PiCodingAgent, 2026-09-17: “i spend all my tokens customizing pi. when everything is perfect, i'll get some work done” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wi77ib/for_what_you_are_using_pi_coding_or_something/paama55/)
  - Complaint, Pi, r/PiCodingAgent, 2026-09-17: “the main question is if "the bloat" actually helps or not. in some benchmarks pi scores way worse than other harnesses which have a "ton of bloat".” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wix9dl/does_anyone_else_keep_track_of_their/paecfsb/)
  - Complaint, Pi, @pidotdev, 2026-09-18: “@pidotdev 4 tools yady yada and still loses against 37 tools xd all oss and runnable <strict_link>” [source](https://twitter.com/837123997993168896/status/2100945849664590207)
  - Complaint, Pi, @pidotdev, 2026-09-20: “@pidotdev need to try if gpt astra is still lazy, refuse to dig deep, try to find ways out in pi after v0.86.0. it seems only me has that problem. astra &lt; 5.6 sol in pi for me. i have to switch to codex to use astra there. astra feels much better in codex than in pi.” [source](https://twitter.com/88011125/status/2101618024893825482)

### What works

- **Extensions you can actually build on** ([MCP servers, plugins, skills and hooks](https://feedbackbench.com/criteria/setup.extensions_mcp.md), [Onboarding, discoverability and documentation](https://feedbackbench.com/criteria/setup.onboarding_docs.md)). Pi's small, open core lets users override built-in tools and ship their own skills, so the harness bends to each workflow.
  Users describe replacing the built-in read, bash, edit and write tools to sandbox them in a container. Others add gated cloud CLI tools and database query helpers.
  
  The open-source angle matters. Users fix what they need on the spot, know exactly what changed, and publish it. One user calls building on the existing foundation 'super convenient'.
  Evidence:
  - Praise, Pi, @pidotdev, 2026-09-11: “@the_judii @pidotdev i prefer it because it is lightweight and minimal. imo, it also tends to get more work done. the ability to extend it by building directly on top of the existing foundation is super convenient.” [source](https://twitter.com/31175473/status/2098474612258754872)
  - Praise, Pi, r/PiCodingAgent, 2026-09-10: “i think the most interesting one i've got is an extension that overrides the built in read, bash, edit and write tools to run them inside a container instead of on the host, which effectively sandboxes it. i've got some smaller extensions that offer it tools for running the aws and kubernetes cli tools (needs to be done on host for credentials), with manual approval needed for each run. i've also got an extension that allows running postgres explain for a database and query, which is useful for optimizing queries.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wcj9bg/who_uses_pi_what_do_you_like_about_it/p8yfopu/)
  - Praise, Pi, @pidotdev, 2026-09-06: “when people ask me why @pidotdev firstly, open source. we improve what we need on the spot and we know exactly what we did. then, we can share it! <strict_link>” [source](https://twitter.com/1399418379522371584/status/2096569419682156958)
  - Praise, Pi, @pidotdev, 2026-09-14: “@pidotdev harness is by far the easiest most intuitive sdk. wiring the harness is so simple and makes stream support a first class principle (finally)” [source](https://twitter.com/1892195856801030144/status/2099574541458759908)

- **Low context use keeps bills small** ([Single prompt, model or effort level consumes disproportionate quota](https://feedbackbench.com/criteria/limits.burn_rate.md), [Prompt cache hits, misses and invalidation](https://feedbackbench.com/criteria/limits.prompt_cache.md), [How much use a plan's price buys](https://feedbackbench.com/criteria/limits.plan_value.md)). A four-tool core and high prompt cache hit rates stretch quotas, so expensive models last far longer on the same plan.
  One user reports an estimated $27 run that really cost $5.72 thanks to cache hits. Another credits Pi's cache rate with keeping a pricey model from draining a weekly limit in half an hour.
  
  Users say it is hard to find an agent using fewer context tokens. Bash as one tool description is cited as where the savings live.
  Evidence:
  - Praise, Pi, r/PiCodingAgent, 2026-09-06: “<strict_link> final estimated cost: $27 real cost: $5,72 !!! hit cache rate: 99,4% !!! i use deepseek api official” [source](https://www.reddit.com/r/PiCodingAgent/comments/1w8ju8c/this_is_why_i_love_pi/p83vv19/)
  - Praise, Pi, @pidotdev, 2026-09-06: “at this point, the only thing keeping astra from burning through my weekly limit in half an hour is @pidotdev's incredible cache hit rate: 99.6% 🚀 also my custom pi extension that show cache ttl countdown timers and ping me via push notification 5m before expiry help a lot 🤓” [source](https://twitter.com/2915475819/status/2096584453376102561)
  - Praise, Pi, @pidotdev, 2026-09-13: “@howaboua i think it's hard to find any agent using less context tokens than @pidotdev 😅 well, amp has some cool tools like painter, librarian, oracle etc. that's why the increased context token size.” [source](https://twitter.com/98821843/status/2099053118654722348)
  - Praise, Pi, @pidotdev, 2026-09-18: “@pidotdev bash buys the whole shell for one tool description. that is where the context saving lives.” [source](https://twitter.com/1549055479875342336/status/2100948485751074862)

- **Any provider, switch mid-task** ([Connecting own API keys, local models and custom endpoints](https://feedbackbench.com/criteria/setup.provider_byok_local.md), [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md), [Using an existing subscription across tools](https://feedbackbench.com/criteria/billing.subscription_portability.md)). Pi is a harness rather than a provider, so users connect their own keys or local models and change models without restarting.
  Users value continuing a task after switching provider or model. They run OpenAI subscriptions, DeepSeek credits, Grok and local models through llama.cpp in one place.
  
  Release posts highlight large catalog refreshes that add new frontier models quickly. Local use works, users say, given a capable GPU and careful quantization choices.
  Evidence:
  - Praise, Pi, @pidotdev, 2026-09-07: “@peilvdog_cn yes. for the same reason i'm currently trying to switch my harness to @pidotdev, so i can switch provider/model and continue the task.” [source](https://twitter.com/305105471/status/2096962219342729511)
  - Praise, Pi, @pidotdev, 2026-09-01: “pi shines as a minimal, open-source harness you fully own and reshape. with just 4 core tools and tiny context overhead, use it to build custom skills/extensions for your exact workflows, switch any model mid-session (including local ones), or run scripts/sdk embeds—complementing the polished but locked-in codex/claude/grok tools. ideal for power users who adapt the agent, not the reverse.” [source](https://twitter.com/1720665183188922368/status/2094838186002313335)
  - Praise, Pi, r/opencode, 2026-09-22: “pi agent is totally free - it is harness, not provider. you can use exactly same models as in opencode harness plus many more as you can connect any provider. in addition, you have about 5000 plugins to enhance/advance your harness” [source](https://www.reddit.com/r/opencode/comments/1wm7pi6/what_the_hell_is_going_on_with_deepseek/pbb5257/)
  - Praise, Pi, r/PiCodingAgent, 2026-09-16: “the best llm model to run with pi currently is qwen-3.8-27b. it needs 16gb of vram (so a high-end gaming gpu). it works very well, you have to find the right balance between model quantization, cache quantization and context size. of course if you expect to run your local model on a standard laptop with a low end gpu then forget it. no local models that run on that kind of hardware is good enough to be really usefull in a coding agent in practice. in a couple of years maybe.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1whunn9/which_subscription_with_pi/pa57cal/)

- **A quiet prompt that keeps models focused** ([Can do the user's kind of task](https://feedbackbench.com/criteria/work.capability.md), [Saving, switching, resuming and rewinding sessions](https://feedbackbench.com/criteria/ui.session_history.md)). The minimal system prompt leaves the model undistracted, which users credit for better results on focused, complex tasks.
  One user runs Pi as a dedicated subagent for document analysis because its sparse prompt beats busier agents. Others say heavy prefill hurts local models, while Pi lets them add only what helps.
  
  Session tooling earns praise too. Users call the /tree session navigator amazing and like how scripted runs land in browsable history.
  Evidence:
  - Praise, Pi, r/PiCodingAgent, 2026-09-12: “i use it as a sub-agent for handling a single complex task that requires focused attention—for example, analyzing complicated documents. pi has a very minimal system prompt that doesn’t distract the model from the task, so the end result is usually much better than when using something like opencode or another coding agent.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wclnuc/how_can_pi_improve_my_daytoday_workflow/p9ce7ue/)
  - Complaint, Pi, r/PiCodingAgent, 2026-09-23: “for external agents omp is fine, for local models like qwen, it dumps a ton of prefill in that actually makes it worse not better. the reason pi works so well it's that you tailor it to your needs and refine it instead of dumping the kitchen sink into you llm as prefill and hoping for the best.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wnvwqa/are_pi_and_ohmypi_are_same/pbkgkvu/)
  - Praise, Pi, @pidotdev, 2026-09-05: “@dhh the best job at orchestration is done by the @pidotdev agent. give it a try, you’ll be surprised by how good it is. and the /tree session navigator is amazing, especially in this context.” [source](https://twitter.com/61703/status/2096249525216120986)
  - Praise, Pi, @pidotdev, 2026-09-20: “@pidotdev asking about which assumptions would simplify the code. started using a skill based on this idea for two weeks. pretty good to cleanup messy initial drafts.” [source](https://twitter.com/510176899/status/2101819155359957186)

### Under the surface

- **Extensions are both the fix and the fault** ([MCP servers, plugins, skills and hooks](https://feedbackbench.com/criteria/setup.extensions_mcp.md), [Updates break working setups](https://feedbackbench.com/criteria/rel.update_breakage.md), [Client crashes, freezes and failed tool execution](https://feedbackbench.com/criteria/rel.client_failures.md)). The ecosystem that fills Pi's gaps also introduces its own breakage, so stability depends on which extensions you stack.
  A long-time user says the only flicker problems came from an extension. Others report version mismatches between Pi and third-party add-ons, and forks shipping same-named extensions without explaining the difference.
  
  Compaction, subagents and display fixes all lean on this layer, so its quality shapes the whole experience.
  Evidence:
  - Praise, Pi, r/PiCodingAgent, 2026-09-14: “in six months i've never had this with pi. the only time i had flicker problems it was an extension i was running.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wfvx2f/any_idea_how_to_resolve_for_this/p9q22uh/)
  - Complaint, Pi, @pidotdev, 2026-09-20: “@pidotdev pi 0.86.0 + pi-antigravity + gemini 并不能一起友好地工作，会报错，至少一个有坑，控制变量怀疑是pi 0.86.0 🤨” [source](https://twitter.com/1666221406441590785/status/2101638507467018713)
  - Praise, Pi, r/PiCodingAgent, 2026-09-03: “i do like the visual component. but i think it should be mandatory for an extension to explain how it differs from predecessors esp. those with the same name.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1w6h2t4/recursive_and_async_claude_code_style_subagents/p7my4fy/)
  - Complaint, Pi, @pidotdev, 2026-09-21: “i built an extension called strudel for @pidotdev which is essentially a gateway for the callable agent primitives. it didn’t really have any performance improvements like i thought but i think jev can fix?” [source](https://twitter.com/1717367940096421888/status/2101918998413676875)

- **Savings depend on the model you pick** ([Prompt cache hits, misses and invalidation](https://feedbackbench.com/criteria/limits.prompt_cache.md), [Single prompt, model or effort level consumes disproportionate quota](https://feedbackbench.com/criteria/limits.burn_rate.md), [How much use a plan's price buys](https://feedbackbench.com/criteria/limits.plan_value.md)). Token efficiency is real, but users credit some cache wins to the provider and still see costly models burn quotas fast.
  One user says near-perfect cache rates come from DeepSeek, not Pi. Others warn that particular models drain usage faster than expected, and that harness-level savings are marginal except on expensive models.
  
  Pi surfaces cache misses, which one user frames as a lesson in context management rather than a fix.
  Evidence:
  - Complaint, Pi, r/PiCodingAgent, 2026-09-07: “99.99% on deepseek harness after 800 millions tokens spent. it's deepseek that is amazing at caching, not so much pi.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1w8ju8c/this_is_why_i_love_pi/p89t9n8/)
  - Complaint, Pi, @pidotdev, 2026-09-27: “it's super interesting to run codex models in the @pidotdev harness vs the codex harness. @badlogicgames has it mention when there's a cache miss, and it's most of them. so if you're upset about how fast your sub gets used up, use it on pi and what it will teach you, may get you to figure out how to manage your context window better.” [source](https://twitter.com/21734113/status/2104223345898319922)
  - Complaint, Pi, r/PiCodingAgent, 2026-09-04: “i use commandcode goat on pi, its probably the most bang for the buck when it comes to using sol, its worked great once i got the commandcode add-on for pi, absolutely don't do the same mistake i did and use luna, it burnt through my usage faster than sol has :d” [source](https://www.reddit.com/r/PiCodingAgent/comments/1w6x3p0/commandcode_vs_chatgpt_plus_for_pi/p7qfrgx/)
  - Complaint, Pi, @pidotdev, 2026-09-18: “@ace_the_agent @pidotdev oui, après franchement, selon le contexte et le modèle, selon moi ce sont des économies de bouts de chandelle qui ne valent pas forcément le coup. sauf sur un modèle cher type astra / fable, où les abonnements se crament à vitesse éclair.” [source](https://twitter.com/1833193917098893313/status/2100894582439551171)

### Fine print

- Most posts come from Pi's own subreddit and X account, which likely skews toward engaged power users.
- Many complaints concern community extensions and forks, making Pi-core issues hard to separate from ecosystem issues.
- Checking, account and most narrow criteria have too few posts to support firm conclusions.

## Top requests

What users ask to add or change, most asked first. 332 author-weeks ask for something. Requests do not change the Feedback Score. Rule: A separate pass by Claude Sonnet 5 reads every counted post and extracts what the author asks the agent or its vendor to add or change, with the criteria it maps to and a short normalised wording; it does not touch the labels or the Feedback Score. Claude Opus 5.5 groups the wordings within each criterion (the first criterion the request maps to) into themes; code counts them. A theme counts distinct author-weeks that ask for it, per agent; across agents, one author-week per agent. Themes asked in fewer than 2 author-weeks, and requests that share no theme, are not shown. Examples: up to 3 posts per theme from different authors, without slurs, preferring posts of 60 to 450 characters, most recent first.

| Rank | Request | Criterion | Author-weeks | Posts |
|---|---|---|---|---|
| 1 | Bring existing subscription into this agent | [Using an existing subscription across tools](https://feedbackbench.com/criteria/billing.subscription_portability.md) | 17 | 17 |
| 2 | Better compaction summary quality and retention | [Context compaction keeps what matters, cheaply and quickly](https://feedbackbench.com/criteria/context.compaction.md) | 6 | 6 |
| 3 | Built-in multi-agent orchestrator mode | [Subagents, parallel agents and orchestrators](https://feedbackbench.com/criteria/work.multi_agent_orchestration.md) | 6 | 6 |
| 4 | Dedicated desktop app | [Install, launch and sign-in](https://feedbackbench.com/criteria/setup.install_signin.md) | 5 | 5 |
| 5 | Live dashboard of subagent status and progress | [Subagents, parallel agents and orchestrators](https://feedbackbench.com/criteria/work.multi_agent_orchestration.md) | 5 | 5 |
| 6 | Multi-provider model choice in one harness | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 5 | 5 |
| 7 | Rewind to checkpoint with code restore | [Saving, switching, resuming and rewinding sessions](https://feedbackbench.com/criteria/ui.session_history.md) | 5 | 5 |
| 8 | Hosted cloud agent execution support | [Cloud and remote sandbox execution](https://feedbackbench.com/criteria/surfaces.cloud_sessions.md) | 4 | 6 |
| 9 | Local model support | [Connecting own API keys, local models and custom endpoints](https://feedbackbench.com/criteria/setup.provider_byok_local.md) | 4 | 4 |
| 10 | Richer extensibility API | [MCP servers, plugins, skills and hooks](https://feedbackbench.com/criteria/setup.extensions_mcp.md) | 4 | 4 |
| 11 | Route only at first turn or manual trigger | [Automatic model routing and fallback](https://feedbackbench.com/criteria/models.routing_auto.md) | 3 | 6 |
| 12 | Agent teams with assignable roles | [Subagents, parallel agents and orchestrators](https://feedbackbench.com/criteria/work.multi_agent_orchestration.md) | 3 | 3 |

### 1. Bring existing subscription into this agent

- Pi, 2026-09-26, r/PiCodingAgent (Reddit): “yeah, i mean, after looking into this, i don't think this is what i need. i actually want to just use pi with opus models directly.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wq1bv2/whats_the_best_option_for_using_anthropic/pc30ait/)
- Pi, 2026-09-25, r/PiCodingAgent (Reddit): “i know there are *many* options out there, but this is still a grey area right? what are you guys using? do you know anyone who has been banned by doing something like this? i have a bunch of friends who use `omp` with their subscriptions and are still kicking. i really want to move from cc to pi, but i am not able to run local models due to my hardware limitations and i would love to start by using my anthropic subscription.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wq1bv2/whats_the_best_option_for_using_anthropic/)
- Pi, 2026-09-24, @pidotdev (X): “i daily drive @pidotdev, but i want to use opus 5.5 for personal use w/ a sub, what to do??? 😭” [source](https://twitter.com/1035016280770785280/status/2102919936679293374)

### 2. Better compaction summary quality and retention

- Pi, 2026-09-27, r/PiCodingAgent (Reddit): “curious as well. i've been reading good things about codex compaction enhancements recently. would be nice to port some of that over to pi if possible.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wrq50c/new_session_handoff_vs_compact_which_do_you_prefer/pcettr9/)
- Pi, 2026-09-22, @pidotdev (X): “@pidotdev remote control native support. better compaction. human-writing like :)” [source](https://twitter.com/89646598/status/2102407468450296138)
- Pi, 2026-09-18, r/PiCodingAgent (Reddit): “extensibility /tree better compaction local state of what happened in the session for analysis / learning. ability to use other models than anthropic. when i do code review on a complex change, it’s so important to have multiple top tier intelligence models run at it to uncover bugs. eg: astra finds bugs that opus wrote ( and vis versa )” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wjtm27/how_to_use_pi_as_a_better_cc/pank9t6/)

### 3. Built-in multi-agent orchestrator mode

- Pi, 2026-09-22, @pidotdev (X): “@pidotdev remaining quota display widget, plannotator review in vscode + herdr orchestration. <strict_link>” [source](https://twitter.com/947783691320905729/status/2102411360743080258)
- Pi, 2026-09-22, @pidotdev (X): “@pidotdev <strict_link> not just a gui solution - but expandeture of capabilities. such as new sub-agent system, browser annotation etc.” [source](https://twitter.com/2059310670131195904/status/2102406216131756254)
- Pi, 2026-09-22, @pidotdev (X): “@pidotdev +1 for first class subagents (not just pi -p in bash, because default pi just freezes there while it delegates the wait, and pi does not do very well with background commands without a lot of manual management) and a better edit tool (minor whitespace errors fail edits)” [source](https://twitter.com/1678448492690419712/status/2102402544693805561)

### 4. Dedicated desktop app

- Pi, 2026-09-27, @pidotdev (X): “genuinely believe a solid desktop app will unlock a crazy amount of new users on to @pidotdev, and excited that pi-gui can play a part in that. it has to be a desktop app focused on pi though, because pi has unique features like /tree and extensions that have to be showcased.” [source](https://twitter.com/1689423238173007873/status/2104240272628404326)
- Pi, 2026-09-25, @pidotdev (X): “@pidotdev i need a official desktop.😭” [source](https://twitter.com/2009214248992559104/status/2103375760665006293)
- Pi, 2026-09-23, @pidotdev (X): “@pidotdev need desktop” [source](https://twitter.com/1992260679697575936/status/2102590528575651956)

### 5. Live dashboard of subagent status and progress

- Pi, 2026-09-25, @pidotdev (X): “@pidotdev also - `pi -p` opens a background subagent. any chance of a foreground subagent? i want a window that opens up on top of my existing pi window, and when i finish it closes itself, drops me back to the previous window with a summary or whatever. call stack :)” [source](https://twitter.com/1678448492690419712/status/2103526499588690092)
- Pi, 2026-09-22, @pidotdev (X): “@pidotdev 我使用做编程时，发现pi给了子代理后并没有任何状态显示，不知道子代理的进度，不知道是不是我的设置不对还是本来就没有这个功能。我让pi自己做了一个子代理的任务进度显示，但是效果不好” [source](https://twitter.com/1816114569456345088/status/2102398739612946722)
- Pi, 2026-09-18, r/PiCodingAgent (Reddit): “seeing the subagent breakdown would be something i'd use. raw token count tells me who was expensive, but i also want to know whether that subagent completed its job or just wandered around for 20 calls. that’s something i look at through braintrust when searching through agent traces, and having a local pi view of it would also be useful.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wj4flu/i_built_pi_session_inspector_to_see_where_my/pakqbs5/)

### 6. Multi-provider model choice in one harness

- Pi, 2026-09-24, r/PiCodingAgent (Reddit): “this looks really useful, also for getting better at prompting and planning as a human :) yes pi support would be great, i use multiple harnesses on the same project” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wp7uqr/would_you_guys_find_this_tool_useful/pbt2y1e/)
- Pi, 2026-09-22, @pidotdev (X): “@pidotdev can we actually run it ? i am dying to run opus in pi since forever. i got codex sub just so i can their models in pi.” [source](https://twitter.com/2685094046/status/2102458590493806948)
- Pi, 2026-09-05, @pidotdev (X): “@mrahmadawais hope the token‑usage metrics in the plan user dashboard could be displayed more clearly, similar to how opencodego presents them. also, could pi add the commandcode plan to the pi‑ai routing? that would make things much more convenient. thanks.@pidotdev” [source](https://twitter.com/1871869440150962176/status/2096062027555025017)

### 7. Rewind to checkpoint with code restore

- Pi, 2026-09-26, @pidotdev (X): “@pidotdev @pidotdev the ability to undo your code changes / updates when you move back with the /timeline feature pls 😃” [source](https://twitter.com/1579853783290580993/status/2103793570155237673)
- Pi, 2026-09-23, @pidotdev (X): “i can't believe no major harness has adopted `/tree` from @pidotdev. some have forks, but man i miss `/tree` so much in codex and @chatgpt. i wish each message had a "rollback" button, which would summarize the tail of the chat up to the selected message.” [source](https://twitter.com/15790969/status/2102812042272874506)
- Pi, 2026-09-09, @pidotdev (X): “@pidotdev can you pls add easy way to restore checkpoint feature to go back in conversation and restart” [source](https://twitter.com/2086150037726539776/status/2097709557414072689)

### 8. Hosted cloud agent execution support

- Pi, 2026-09-23, @pidotdev (X): “@pidotdev @mattlam_ @badlogicgames please, make this happen!” [source](https://twitter.com/568738762/status/2102808779557675038)
- Pi, 2026-09-23, @pidotdev (X): “@samuelvrablik @pidotdev @badlogicgames ye it's easy to build one for myself, but i want the same ux as cursor's where anyone can spin up cloud agents easily, adn the infra is managed for you. also many different cloud agents, that's the goal.” [source](https://twitter.com/1689423238173007873/status/2102796386106487164)
- Pi, 2026-09-23, @pidotdev (X): “@pidotdev something like opencode go will be highly appreciated.” [source](https://twitter.com/17648829/status/2102714998858600756)

### 9. Local model support

- Pi, 2026-09-23, @pidotdev (X): “@pidotdev please make initial setup easier for vllm/sglang/llama.cpp! writing models.json by hand is a pain and the /login path for llama.cpp doesn't seem to work for me (and i end up using nano for models.json like a pleb)” [source](https://twitter.com/50067714/status/2102622692637897183)
- Pi, 2026-09-11, r/PiCodingAgent (Reddit): “similar question, the ui looks sleek, but looks it requires to select a supported provider. i run pi on local llm, so can’t use this?” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wcp3b7/supernova_a_minimal_opinionated_and_sleek/p93aa59/)
- Pi, 2026-09-17, r/PiCodingAgent (Reddit): “this'll be great when there's a true local version.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wimfhg/piwarden_a_jevpowered_second_pair_of_eyes_for_pi/pacrvb8/)

### 10. Richer extensibility API

- Pi, 2026-09-27, r/PiCodingAgent (Reddit): “giga bloat even with system prompt off no way to trim down tool output or set limits unless u waste a ton of tokens making a wrapper no custom extensions etc” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wrjr75/the_good_opensource_harness/pcgm7ya/)
- Pi, 2026-09-22, @pidotdev (X): “@pidotdev let me override more of `pi` 👀 i want to be able to take ownership of the message parsing / handling directly at the response or websocket layer. let me own the session storage. let me change how messages are ordered.” [source](https://twitter.com/849247899770925057/status/2102420812028633454)
- Pi, 2026-09-22, @pidotdev (X): “@pidotdev just keep it simple and make it so adding new functionality is also dead simple more api access to the internals. tried writing ohmypi time travel rules and pi doesn't allow 100% reproduction” [source](https://twitter.com/18738053/status/2102406817233989777)

### 11. Route only at first turn or manual trigger

- Pi, 2026-09-20, r/PiCodingAgent (Reddit): “of course i know that. that's why you'd want the power to manually switch, and to signal when a manual switch might be prudent.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wlczow/i_think_i_found_the_best_use_case_for_jev_and_pi/paz5xwx/)
- Pi, 2026-09-20, r/PiCodingAgent (Reddit): “> skills hook. after each user message, determine which new skills should be loaded. this would replace current skill functionality. you could have a huge skill index without polluting llm context. this sounds like a good use of the local laya model. or some sort of tools suggestion at the very least just a suggestion engine (i don't want anything routing or breaking context without my perm)” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wlczow/i_think_i_found_the_best_use_case_for_jev_and_pi/payss4m/)

### 12. Agent teams with assignable roles

- Pi, 2026-09-25, r/PiCodingAgent (Reddit): “use it once and given up. it can not work like build up an agent teams and let different session collobration together.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wogo23/has_anyone_tried_piintercom/pbxuqz1/)
- Pi, 2026-09-22, @pidotdev (X): “@pidotdev @indydevdan composable ai developer workflows with custom pi agents. maybe some "primitive" agents that can be built into pi agent teams, particularly for sdlc workflows” [source](https://twitter.com/2758152969/status/2102415098987888922)
- Pi, 2026-09-03, r/PiCodingAgent (Reddit): “i'm struggling to get started with pi. today i use omp, it's really great, but i don't know exactly what i need and my needings, but deffinitely i need the subagents and model roles.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1w593fd/i_benchmarked_pi_against_its_own_fork_omp_more/p7jr8ql/)

## Facts

| Fact | Value |
|---|---|
| Version | v0.84.x (Aug 2026); npm @earendil-works/pi-coding-agent |
| Released | First release: 2025-08. Joined Earendil: 2026-04-08 |
| Price | Free (MIT). Model usage billed by the chosen provider's API or subscription |
| Model | Multi-provider, bring-your-own-key or subscription |
| Surface | CLI (terminal) |

## Sources

| Channel | Source | Posts |
|---|---|---|
| Reddit | r/PiCodingAgent | 4114 |
| X | @pidotdev | 1973 |
| Reddit | Posts that name it | 126 |

## Better than peers on

How much use a plan's price buys, Single prompt, model or effort level consumes disproportionate quota, Prompt cache hits, misses and invalidation, Using an existing subscription across tools, Connecting own API keys, local models and custom endpoints, MCP servers, plugins, skills and hooks, Onboarding, discoverability and documentation, Which models are offered on a plan and when, Context compaction keeps what matters, cheaply and quickly, Can do the user's kind of task, How the interface shows work, and what the user can configure, Saving, switching, resuming and rewinding sessions, Latency, throughput and fast mode, Client crashes, freezes and failed tool execution

## Worse than peers on

None.

## All 63 criteria

Criterion love: 0.5 is the category norm. n: rated author-weeks.

### Paying and limits: Better than peers (customer love 0.686, n 245)

| Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|
| [Single prompt, model or effort level consumes disproportionate quota](https://feedbackbench.com/criteria/limits.burn_rate.md) | Better than peers | 0.646 | 0.605–0.681 | 84 | 43 | 41 |
| [Prompt cache hits, misses and invalidation](https://feedbackbench.com/criteria/limits.prompt_cache.md) | Better than peers | 0.536 | 0.510–0.565 | 61 | 39 | 22 |
| [Using an existing subscription across tools](https://feedbackbench.com/criteria/billing.subscription_portability.md) | Better than peers | 0.535 | 0.509–0.558 | 45 | 26 | 19 |
| [How much use a plan's price buys](https://feedbackbench.com/criteria/limits.plan_value.md) | Better than peers | 0.567 | 0.539–0.593 | 42 | 32 | 10 |
| [Usage meter visibility and accuracy](https://feedbackbench.com/criteria/limits.usage_meter.md) | Too few posts | 0.558 | 0.515–0.606 | 11 | 6 | 5 |
| [Free tier and free model availability and limits](https://feedbackbench.com/criteria/billing.free_tier.md) | Too few posts | 0.504 | 0.491–0.517 | 8 | 5 | 3 |
| [Short rolling usage window blocks or interrupts work](https://feedbackbench.com/criteria/limits.window_interrupts_work.md) | Too few posts | 0.491 | 0.484–0.497 | 6 | 0 | 6 |
| [Pricing and plan terms stated clearly and consistently](https://feedbackbench.com/criteria/billing.pricing_clarity.md) | Too few posts | 0.495 | 0.489–0.499 | 4 | 0 | 4 |
| [Quota reset timing and bonus or banked resets](https://feedbackbench.com/criteria/limits.reset_schedule.md) | Too few posts | 0.495 | 0.490–0.500 | 3 | 0 | 3 |
| [Price, allowance or plan terms changed](https://feedbackbench.com/criteria/limits.allowance_change.md) | Too few posts | 0.499 | 0.496–0.500 | 1 | 0 | 1 |
| [Pay-as-you-go overage, fallback billing and spend caps](https://feedbackbench.com/criteria/billing.overage_charges.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 |

Most recent posts:

- Praise, 2026-09-26, r/PiCodingAgent (Reddit): “because you can use your max subscription” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wqmcmk/poll_claude_sub_with_pi_what_method_do_you_use/pc5bdn3/)
- Praise, 2026-09-26, r/LocalLLaMA (Reddit): “can you actually remove / override the heavy system prompt though? according to my searching (and asking llms) you can only *add* to the system prompt. also, even when you enable only the bare minimum tool calling like file read / write and a couple others it's still like 8k tokens for even just a 'hello world' style prompt. some of us can run models relatively smoothly on minimal hardware with pr…” [source](https://www.reddit.com/r/LocalLLaMA/comments/1wq9ivr/what_ide_to_use_for_local_models/pc3lhdp/)
- Praise, 2026-09-25, r/PiCodingAgent (Reddit): “i’m fairly suspicious of anthropic being able to detect most integrations that allow claude in the drivers seat on pi. i have too much important shit on chat to treat it like a burner account. i personally only do claude subagents on pi via headless since anthropic somewhat sanctions that behavior. if i need claude to drive, i just use claude code for now. pi will proxy to claude code just fine.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wq1bv2/whats_the_best_option_for_using_anthropic/pc0bxyh/)
- Complaint, 2026-09-27, r/PiCodingAgent (Reddit): “every use case is different. personally, i prefer a lean approach where the context window stays small and gets summarized regularly, while solid documentation lets me spin up new sessions with fast project context recovery. - i avoid large context windows because i don't have much vram. once a task is done, i compress the context using my `pi-refine-compact` extension, so the llm stays up to spee…” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wrq50c/new_session_handoff_vs_compact_which_do_you_prefer/pcf5zei/)
- Complaint, 2026-09-27, r/PiCodingAgent (Reddit): “i was able to use it for a little while before i started hitting request limits, and it’s never really reset. now i just have a qwen model that fits in my vram running through ollama that i use in pi (using ollama as the provider).” [source](https://www.reddit.com/r/PiCodingAgent/comments/1vjurqz/considering_claude_code_pi_worth_it/pcf7g1b/)
- Complaint, 2026-09-27, r/PiCodingAgent (Reddit): “if i had an anthropic plan, i wouldn't risk it - dont they say they can ban you for this? how do you use /tree and whay extensions do you use?” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wrjr75/the_good_opensource_harness/pcgsa4u/)

### Setting up and connecting: Better than peers (customer love 0.632, n 307)

| Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|
| [MCP servers, plugins, skills and hooks](https://feedbackbench.com/criteria/setup.extensions_mcp.md) | Better than peers | 0.586 | 0.558–0.619 | 201 | 136 | 65 |
| [Connecting own API keys, local models and custom endpoints](https://feedbackbench.com/criteria/setup.provider_byok_local.md) | Better than peers | 0.532 | 0.505–0.557 | 60 | 41 | 19 |
| [Onboarding, discoverability and documentation](https://feedbackbench.com/criteria/setup.onboarding_docs.md) | Better than peers | 0.549 | 0.513–0.583 | 36 | 15 | 21 |
| [Install, launch and sign-in](https://feedbackbench.com/criteria/setup.install_signin.md) | Too few posts | 0.512 | 0.486–0.540 | 21 | 6 | 15 |
| [IDE and editor integration](https://feedbackbench.com/criteria/setup.ide_integration.md) | Too few posts | 0.512 | 0.499–0.526 | 8 | 6 | 2 |

Most recent posts:

- Praise, 2026-09-27, r/PiCodingAgent (Reddit): “pi for my use case. i also use opencode for free models but pi when i use my own api” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wrjr75/the_good_opensource_harness/pcd80nj/)
- Praise, 2026-09-27, r/PiCodingAgent (Reddit): “pi has extension, i've built my own for my harness too. imho it's a way to go instead of compactions” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wrq50c/new_session_handoff_vs_compact_which_do_you_prefer/pceoo4i/)
- Praise, 2026-09-27, r/PiCodingAgent (Reddit): “i'm using pi-blackhole and never had problems with it. what did you used?” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wrq50c/new_session_handoff_vs_compact_which_do_you_prefer/pcezess/)
- Complaint, 2026-09-27, r/PiCodingAgent (Reddit): “ok, honestly i do not understand why this is popular, this is literally just vibed code knock off from codex/ux, 'make an app that has pi backend, but look exactly like codex', done, one prompt, second, for any modern harness/app, esp claude code, one plugin, you could run anything, configure in anyways you want, codex is less customizable for obvious reasons, i cant believe how ignorant people ar…” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wcp3b7/supernova_a_minimal_opinionated_and_sleek/pcb85p7/)
- Complaint, 2026-09-27, r/PiCodingAgent (Reddit): “giga bloat even with system prompt off no way to trim down tool output or set limits unless u waste a ton of tokens making a wrapper no custom extensions etc” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wrjr75/the_good_opensource_harness/pcgm7ya/)
- Complaint, 2026-09-27, @pidotdev (X): “@pigcodingagent @pidotdev i use opencode in pi but i tried and couldn't in pig. it would be great if you added opencode as a login option.” [source](https://twitter.com/299687169/status/2104021066016580085)

### Choosing models: Better than peers (customer love 0.581, n 79)

| Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|
| [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | Better than peers | 0.556 | 0.523–0.586 | 36 | 22 | 14 |
| [Automatic model routing and fallback](https://feedbackbench.com/criteria/models.routing_auto.md) | Too few posts | 0.547 | 0.517–0.579 | 29 | 16 | 13 |
| [Reasoning effort setting and its defaults](https://feedbackbench.com/criteria/models.effort_control.md) | Too few posts | 0.502 | 0.488–0.518 | 10 | 5 | 5 |
| [Quality got worse or better over time](https://feedbackbench.com/criteria/models.quality_drift.md) | Too few posts | 0.490 | 0.482–0.497 | 6 | 0 | 6 |

Most recent posts:

- Praise, 2026-09-27, @pidotdev (X): “opus 5.5 completely dominates gpt-6 sol at blender 3d pelican 🦩 prompt: animate a looping 3d pelican on a bicycle in blender and opus 5.5 came back with the more charming ride use opus 5.5 in @pidotdev 👉 <strict_link> <strict_link>” [source](https://twitter.com/2001569273681186823/status/2104200810293014713)
- Praise, 2026-09-25, r/PiCodingAgent (Reddit): “the point everyone is making is that luna max can function equivalently to the larger models and thus should be considered for the same comparison. definitely less thinking is faster and works well with some guidance” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wpcis6/opus_55_vs_gpt6_sol_luna_in_piagent_results_on_my/pbvj728/)
- Praise, 2026-09-25, r/PiCodingAgent (Reddit): “i love to see deepseek flash here, its my favorite alternative to 5.6luna” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wpcis6/opus_55_vs_gpt6_sol_luna_in_piagent_results_on_my/pbwccta/)
- Complaint, 2026-09-27, r/PiCodingAgent (Reddit): “had too. most of the time it would hit the context window every single time and then didn't answer the prompt. also didn't see much difference between thinking modes, but that might just be my perception after 30m of waiting for the model to actually come to a conclusion. any conclusion at all.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wnmsho/i_want_to_believe_in_local_llms_for_coding_but/pc9s2g3/)
- Complaint, 2026-09-26, r/PiCodingAgent (Reddit): “reasoning off for all requests? that's crazy for a model that is designed for massive thinking traces.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wnmsho/i_want_to_believe_in_local_llms_for_coding_but/pc4dne9/)
- Complaint, 2026-09-26, @pidotdev (X): “@shantanugoel @pidotdev its not a good model, just use luna6” [source](https://twitter.com/1448626313619705856/status/2103773838320550066)

### Instructing and context: Typical (customer love 0.529, n 179)

| Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|
| [Context compaction keeps what matters, cheaply and quickly](https://feedbackbench.com/criteria/context.compaction.md) | Better than peers | 0.536 | 0.504–0.567 | 73 | 34 | 39 |
| [Memory and state carried across sessions](https://feedbackbench.com/criteria/context.session_memory.md) | Typical | 0.508 | 0.482–0.535 | 38 | 20 | 18 |
| [Persistent project rules files are read and obeyed](https://feedbackbench.com/criteria/context.instruction_files.md) | Too few posts | 0.501 | 0.482–0.519 | 19 | 10 | 9 |
| [Direct in-prompt instructions and caps are followed](https://feedbackbench.com/criteria/context.instruction_following.md) | Too few posts | 0.496 | 0.476–0.518 | 17 | 4 | 13 |
| [Finding the right files in the codebase](https://feedbackbench.com/criteria/context.codebase_retrieval.md) | Too few posts | 0.494 | 0.476–0.511 | 17 | 6 | 11 |
| [Output degrades as the context window fills](https://feedbackbench.com/criteria/context.long_context_decay.md) | Too few posts | 0.497 | 0.478–0.523 | 15 | 2 | 13 |
| [Images, PDFs and file attachments as input](https://feedbackbench.com/criteria/context.attachments.md) | Too few posts | 0.502 | 0.490–0.516 | 9 | 4 | 5 |
| [Asks the user versus guessing](https://feedbackbench.com/criteria/context.clarifying_questions.md) | Too few posts | 0.505 | 0.490–0.520 | 8 | 4 | 4 |

Most recent posts:

- Praise, 2026-09-27, r/PiCodingAgent (Reddit): “i use observational memory and forgot about compaction.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wrq50c/new_session_handoff_vs_compact_which_do_you_prefer/pcen0k3/)
- Praise, 2026-09-27, r/PiCodingAgent (Reddit): “yeah, i should've been using this since yesterday. i built a summarization workflow myself but this is actually better. thanks again” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wrq50c/new_session_handoff_vs_compact_which_do_you_prefer/pcfg2ad/)
- Praise, 2026-09-27, r/PiCodingAgent (Reddit): “yeah so like i said, the tool restriction accomplishes what i wanted. this is a follow up post on how to best go about that piece. almost nothing to do with your comment, which is also unhelpful as tool restriction is much more effective than just tweaking prompt .md files.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wrte31/alternative_to_tool_profiles_for_better_subagent/pcgoayr/)
- Complaint, 2026-09-27, r/PiCodingAgent (Reddit): “had too. most of the time it would hit the context window every single time and then didn't answer the prompt. also didn't see much difference between thinking modes, but that might just be my perception after 30m of waiting for the model to actually come to a conclusion. any conclusion at all.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wnmsho/i_want_to_believe_in_local_llms_for_coding_but/pc9s2g3/)
- Complaint, 2026-09-27, r/PiCodingAgent (Reddit): “check out cortexkit i have been running their complete suite for the past week and i forgot about the concept of context and compaction. i have a nice workflow built up around and preceding the cortexkit suite that provides durable context but.. i will say this, after this week trial i have it i am keeping it on both machines aft => replaces pis 4 tools with its own + 3 more that all hinge in the…” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wrq50c/new_session_handoff_vs_compact_which_do_you_prefer/pcepylh/)
- Complaint, 2026-09-27, r/PiCodingAgent (Reddit): “and you don't care that your input token size explodes if you never compact? that would suck claude usage with a boosted straw” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wrq50c/new_session_handoff_vs_compact_which_do_you_prefer/pcesnfy/)

### Doing the work: Better than peers (customer love 0.554, n 328)

| Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|
| [Can do the user's kind of task](https://feedbackbench.com/criteria/work.capability.md) | Better than peers | 0.544 | 0.511–0.578 | 163 | 116 | 47 |
| [Subagents, parallel agents and orchestrators](https://feedbackbench.com/criteria/work.multi_agent_orchestration.md) | Typical | 0.488 | 0.455–0.520 | 88 | 50 | 38 |
| [Spins, loops or gets stuck without progress](https://feedbackbench.com/criteria/work.stuck_loops.md) | Too few posts | 0.567 | 0.499–0.626 | 26 | 5 | 21 |
| [Risky or irreversible actions without confirmation](https://feedbackbench.com/criteria/work.destructive_actions.md) | Too few posts | 0.516 | 0.488–0.544 | 20 | 6 | 14 |
| [Long unattended runs and goal/loop mode](https://feedbackbench.com/criteria/work.long_running_autonomy.md) | Too few posts | 0.508 | 0.488–0.525 | 14 | 12 | 2 |
| [Tool approval prompts and autonomy modes](https://feedbackbench.com/criteria/work.permission_prompts.md) | Too few posts | 0.540 | 0.515–0.567 | 13 | 10 | 3 |
| [Does unrequested work or over-engineers](https://feedbackbench.com/criteria/work.scope_overreach.md) | Too few posts | 0.484 | 0.475–0.492 | 12 | 0 | 12 |
| [Plan-before-edit mode](https://feedbackbench.com/criteria/work.plan_mode.md) | Too few posts | 0.498 | 0.481–0.513 | 12 | 5 | 7 |
| [Computer use and browser control](https://feedbackbench.com/criteria/work.computer_browser_use.md) | Too few posts | 0.491 | 0.477–0.502 | 7 | 2 | 5 |
| [Breaks existing code or reintroduces bugs](https://feedbackbench.com/criteria/work.regressions_introduced.md) | Too few posts | 0.500 | 0.489–0.514 | 5 | 1 | 4 |
| [Frontend and visual UI output](https://feedbackbench.com/criteria/work.frontend_ui.md) | Too few posts | 0.503 | 0.496–0.511 | 4 | 3 | 1 |
| [Length and clarity of replies, summaries and comments](https://feedbackbench.com/criteria/work.response_verbosity.md) | Too few posts | 0.504 | 0.493–0.519 | 4 | 2 | 2 |
| [Safety filters block legitimate coding tasks](https://feedbackbench.com/criteria/work.safety_refusals.md) | Too few posts | 0.496 | 0.490–0.500 | 3 | 0 | 3 |
| [Git commits, branches and sync](https://feedbackbench.com/criteria/work.git_workflow.md) | Too few posts | 0.496 | 0.491–0.500 | 2 | 0 | 2 |
| [Diagnosing and fixing reported bugs](https://feedbackbench.com/criteria/work.bug_diagnosis.md) | Too few posts | 0.502 | 0.500–0.506 | 1 | 1 | 0 |
| [Stops mid-task or answers instead of acting](https://feedbackbench.com/criteria/work.premature_stop.md) | Too few posts | 0.498 | 0.494–0.500 | 1 | 0 | 1 |
| [Games checks instead of fixing the problem](https://feedbackbench.com/criteria/work.reward_hacking.md) | Too few posts | 0.499 | 0.496–0.500 | 1 | 0 | 1 |
| [Caves to or argues with the user's judgement](https://feedbackbench.com/criteria/work.sycophancy_pushback.md) | Too few posts | 0.508 | 0.500–0.524 | 1 | 1 | 0 |

Most recent posts:

- Praise, 2026-09-27, r/PiCodingAgent (Reddit): “my orchestrator does nothing except for delegating and passing messages. i've had builds running for almost 24 hours and the orchestrator at the end is still below 200k context. for long builds it's super efficient.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wo6lr8/how_to_better_enforce_subagent_delegation/pcatr9y/)
- Praise, 2026-09-27, r/PiCodingAgent (Reddit): “yes i know the system prompt is big. you said nothing about capabilities. have you actually tried pi with all the added features? things like subagents, memory, worktrees, and many others are all built in, and you need to add these to pi. try to be objective and compare the same things” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wrjr75/the_good_opensource_harness/pcgpxe3/)
- Praise, 2026-09-27, r/PiCodingAgent (Reddit): “i use pi normaly i was trying out opus 5.5 today on claude code and it feels awful vs pi with /tree and all my custom stuff now im looking at the sdk claude pi extension, having model edit and reduce token bloat, hoping it won't get me banned” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wrjr75/the_good_opensource_harness/pcgqytv/)
- Complaint, 2026-09-27, r/PiCodingAgent (Reddit): “the codebase is very well documented, even using a code mapper to save on the scanning, the prompts were crystal clear, with the right context being provided, and still it went off and reasoned about it for a huge amount of time. one of the tests were made in little coder and the harness even tried to tell the model to stop thinking and implement as it already had the solution, but to no avail, it…” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wnmsho/i_want_to_believe_in_local_llms_for_coding_but/pcbw9u1/)
- Complaint, 2026-09-27, r/PiCodingAgent (Reddit): “yeah. on all my tests accuracy was quite shit, and jev was consistently outperformed by glm 5.3 flash, for anything requiring decision making. but hey, it's fast 🤷♂️” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wnl7da/pi_can_now_use_jev_and_more/pce99cd/)
- Complaint, 2026-09-27, r/PiCodingAgent (Reddit): “pretty much it. even if the system prompt says they have to delegate, it usually tries and fails once in order to realize it really has to delegate. i guess you could instruct it to use a classifier model to determine if it needs to delegate and to which agent.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wrte31/alternative_to_tool_profiles_for_better_subagent/pcfl1l2/)

### Checking and finishing: Too few posts (customer love 0.520, n 15)

| Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|
| [Agent-performed code review finds real issues](https://feedbackbench.com/criteria/verify.agent_code_review.md) | Too few posts | 0.512 | 0.504–0.521 | 7 | 7 | 0 |
| [Reviewing and approving the agent's changes](https://feedbackbench.com/criteria/verify.change_review_ui.md) | Too few posts | 0.497 | 0.489–0.506 | 4 | 1 | 3 |
| [Claims work is done or fixed when it is not](https://feedbackbench.com/criteria/verify.false_completion.md) | Too few posts | 0.516 | 0.496–0.557 | 2 | 1 | 1 |
| [Builds, tests or runs its own changes](https://feedbackbench.com/criteria/verify.self_testing.md) | Too few posts | 0.505 | 0.500–0.513 | 2 | 2 | 0 |

Most recent posts:

- Praise, 2026-09-26, r/PiCodingAgent (Reddit): “also nice edited comments haha, sad to see the worlds sharpest dev not be able to help us. and you can go look at the repository tests or maybe live demo attached before saying no testing is happening lol. would love to see some of your work oh great one. anywho thanks again for superb feedback, i'll go put some of the first effort ever in fixing this so we don't let another one of you extremely v…” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wqe2v1/meet_pig_the_pi_coding_harness_that_is_yours_but/pc9g9jc/)
- Praise, 2026-09-25, r/PiCodingAgent (Reddit): “yes in my experience, its a work horse on medium and also good at reviews, catching multiple bugs in both astra and sols work” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wpcis6/opus_55_vs_gpt6_sol_luna_in_piagent_results_on_my/pbze4yj/)
- Praise, 2026-09-23, @pidotdev (X): “@miguelriosen @pidotdev keeping the workflow as a dsl file means it shows up in code review as a diff, which a drag-and-drop canvas never gives you.” [source](https://twitter.com/2058824892238209024/status/2102900988915093764)
- Complaint, 2026-09-22, @pidotdev (X): “@pidotdev the number of extension with "diff"/"review" in title or description. that's a clear signal to improve diff.” [source](https://twitter.com/618819434/status/2102497902077603913)
- Complaint, 2026-09-09, r/PiCodingAgent (Reddit): “ok looks cool but... what value does it actually bring check a session change visually? not a bad concept, but imho an entire application for a functionallity that pi users barely check (yes, i'm pointing to loop/harness users) won't bring so much help. maybe a smaller case as plugin for vs code or obsidian may make more sense” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wbef1h/i_opened_a_pi_session_as_an_editable_map/p8q2d78/)
- Complaint, 2026-09-04, @pidotdev (X): “the useful line is in the log, not the prompt. v7 already retired pi+grok for slower runs and fabricated verification. switching because of weekly limits doesn’t fix that. lock the harness contract first: what “done” means, how you verify, what you do when the model lies. then change the model.” [source](https://twitter.com/1801539591427543040/status/2095871036294119875)

### Interface and sessions: Better than peers (customer love 0.564, n 171)

| Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|
| [How the interface shows work, and what the user can configure](https://feedbackbench.com/criteria/ui.display_settings.md) | Better than peers | 0.553 | 0.517–0.589 | 115 | 55 | 60 |
| [Saving, switching, resuming and rewinding sessions](https://feedbackbench.com/criteria/ui.session_history.md) | Better than peers | 0.584 | 0.556–0.613 | 35 | 28 | 7 |
| [Mobile, remote-control and voice access](https://feedbackbench.com/criteria/surfaces.remote_mobile.md) | Too few posts | 0.506 | 0.489–0.522 | 15 | 9 | 6 |
| [Cloud and remote sandbox execution](https://feedbackbench.com/criteria/surfaces.cloud_sessions.md) | Too few posts | 0.498 | 0.482–0.510 | 7 | 4 | 3 |
| [Stopping and steering a running agent](https://feedbackbench.com/criteria/ui.interrupt_steer.md) | Too few posts | 0.503 | 0.494–0.513 | 5 | 3 | 2 |

Most recent posts:

- Praise, 2026-09-26, r/PiCodingAgent (Reddit): “i'm a bit rusty but i am loving the interface” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wqe2v1/meet_pig_the_pi_coding_harness_that_is_yours_but/pc4od15/)
- Praise, 2026-09-25, r/PiCodingAgent (Reddit): “being able to check on coding stuff from your phone is actually pretty handy. nice little addition” [source](https://www.reddit.com/r/PiCodingAgent/comments/1w29fyl/open_sourced_the_mobile_client_i_built_for_my/pbye1pe/)
- Praise, 2026-09-25, r/PiCodingAgent (Reddit): “i ended up building my own ide (using monaco) that integrates pi that runs entirely as a web app so i can access it form anywhere, it runs on a vps, if i want to work on something i just clone the repo on the vm and send my prompt, i barely even use my local machine anymore.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wpxese/how_do_you_handle_agentic_dev_across_multiple/pc0h8k0/)
- Complaint, 2026-09-27, r/PiCodingAgent (Reddit): “giga bloat even with system prompt off no way to trim down tool output or set limits unless u waste a ton of tokens making a wrapper no custom extensions etc” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wrjr75/the_good_opensource_harness/pcgm7ya/)
- Complaint, 2026-09-25, r/PiCodingAgent (Reddit): “it's pretty but i don't see how it would be useful. i keep all the history i need to retain in my git history .” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wp7uqr/would_you_guys_find_this_tool_useful/pbxfoql/)
- Complaint, 2026-09-25, r/PiCodingAgent (Reddit): “i want my agent to be able to give me a clickable link to find a folder (like "file:/// ...") but even though they respond to mouseover and a tool tip comes up saying "ctrl+click to follow link" they don't work. i am using powershell 7 with pi running in it. my agent had some long complicated idea about how to fix this but does anyone know if i am just overlooking a powershell setting or if this i…” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wpuic2/clickable_links_in_chat_possible/)

### Reliability and speed: Better than peers (customer love 0.592, n 119)

| Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|
| [Latency, throughput and fast mode](https://feedbackbench.com/criteria/rel.response_speed.md) | Better than peers | 0.542 | 0.514–0.568 | 54 | 31 | 23 |
| [Client crashes, freezes and failed tool execution](https://feedbackbench.com/criteria/rel.client_failures.md) | Better than peers | 0.580 | 0.519–0.633 | 46 | 10 | 36 |
| [Updates break working setups](https://feedbackbench.com/criteria/rel.update_breakage.md) | Too few posts | 0.516 | 0.490–0.546 | 14 | 4 | 10 |
| [Outages, server errors and capacity or rate errors](https://feedbackbench.com/criteria/rel.service_errors.md) | Too few posts | 0.521 | 0.484–0.566 | 11 | 2 | 9 |

Most recent posts:

- Praise, 2026-09-27, r/PiCodingAgent (Reddit): “yeah. on all my tests accuracy was quite shit, and jev was consistently outperformed by glm 5.3 flash, for anything requiring decision making. but hey, it's fast 🤷♂️” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wnl7da/pi_can_now_use_jev_and_more/pce99cd/)
- Praise, 2026-09-27, r/PiCodingAgent (Reddit): “pretty good, these days my pi is stock plus one extension showing context usage in detail and one custom footer override to apply custom styling. i do maintain one patch for the llama.cpp provider to enable per model and session thinking levels. performance is fast but i use a dual radeon r9700 setup. i pretty much work offline and with no delays.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1vjurqz/considering_claude_code_pi_worth_it/pch2ftv/)
- Praise, 2026-09-26, r/PiCodingAgent (Reddit): “nice work. i tried using it. it is fast. i tried using few extensions but they don’t seem to work with pig. mcp-adapter, pi-web-access, pi-rtk-optimizer” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wqe2v1/meet_pig_the_pi_coding_harness_that_is_yours_but/pc3o4sz/)
- Complaint, 2026-09-26, r/PiCodingAgent (Reddit): “oof thank you so much definitely was a regression.... hot fix coming. messed up a merge conflict resolve on launch 🤡” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wqe2v1/meet_pig_the_pi_coding_harness_that_is_yours_but/pc3qaf0/)
- Complaint, 2026-09-26, r/PiCodingAgent (Reddit): “every time i turn on my mac and see so many node processes, it’s really terrible. i think i‘ll try.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wqe2v1/meet_pig_the_pi_coding_harness_that_is_yours_but/pc4svol/)
- Complaint, 2026-09-26, r/PiCodingAgent (Reddit): “would be nice if it worked. reality is that it doesn't tried it with stock settings coloring is bugged out, stock \`read\` tool is failing with jsonschema validation so it's just yet another sloppily coded port from one language to another with no real effort put into the maintenance apart from tons of tokens from the company leeched” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wqe2v1/meet_pig_the_pi_coding_harness_that_is_yours_but/pc9bfwt/)

### Account and support: Typical (customer love 0.512, n 30)

| Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|
| [Account bans and access restrictions](https://feedbackbench.com/criteria/account.bans_restrictions.md) | Too few posts | 0.508 | 0.474–0.556 | 20 | 2 | 18 |
| [Data retention, training use and deployment isolation](https://feedbackbench.com/criteria/account.data_privacy.md) | Too few posts | 0.508 | 0.490–0.529 | 7 | 2 | 5 |
| [Support, refunds and issue handling](https://feedbackbench.com/criteria/account.support.md) | Too few posts | 0.505 | 0.494–0.521 | 3 | 1 | 2 |
| [Wrong charges, failed payments and plan provisioning](https://feedbackbench.com/criteria/account.billing_errors.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 |

Most recent posts:

- Praise, 2026-09-25, r/PiCodingAgent (Reddit): “i use pi-web and have used claude x20 max plan for over 8 months, no issues at all.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1whunn9/which_subscription_with_pi/pc2ga42/)
- Praise, 2026-09-25, r/PiCodingAgent (Reddit): “hello, i made pi-claude-request-compat ([<strict_link>). instead of creating another provider, it reuses pi’s existing anthropic login, models, and streaming, then adds a claude code compatibility layer to outgoing requests. [<strict_link> i’m using a new github account for privacy, so the code, automated security scans, dependency audits, and signed release provenance are public for anyone to che…” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wpyjrs/yet_another_pi_addon_claude_code_compatibility/)
- Praise, 2026-09-23, r/PiCodingAgent (Reddit): “i don’t know what you consider a “proper gpu” but mine can handle multiple requests with 32k+ contexts. i value local/privacy over convenience especially since every call to anthropic or openai motivates them to buy more hardware which raises prices for regular consumers wanting to game, or do ai locally. i’ll just stick to pi-subagents until your extension proves itself. clearly you’re not focuse…” [source](https://www.reddit.com/r/PiCodingAgent/comments/1woif1b/i_built_pi_herdsman_for_async_subagents_with/pbnszlp/)
- Complaint, 2026-09-26, r/PiCodingAgent (Reddit): “using that at the moment. haven't been banned yet” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wq1bv2/whats_the_best_option_for_using_anthropic/pc2ufcj/)
- Complaint, 2026-09-26, r/PiCodingAgent (Reddit): “best way... just got my secondary account banned using omp. guys, tos are rigid, does not risk your main accounts.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wqmcmk/poll_claude_sub_with_pi_what_method_do_you_use/pc6fpzc/)
- Complaint, 2026-09-25, r/PiCodingAgent (Reddit): “yep, it’s closed source, i’m comparing the requests the installed cli constructs, not using its source code. the addon reproduces the observed headers, metadata, billing/checksum fields, and tool naming. that said, “completely indistinguishable” was overstating it. matching those fields doesn’t prove anthropic can’t distinguish the clients or enforce restrictions. the account risk is still real.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wpyjrs/yet_another_pi_addon_claude_code_compatibility/pc07p64/)
