有人真的用过 mini swe agent 来 debug 或是开发吗

发布于

之前看 Deep Swe 说 [mini swe agent](https://github.com/swe-agent/mini-swe-agent) 在 debug 中的 token 消耗和成功率都比 Codex 和 Claude Code 更好。我今天自己跑了一下,使用 GPT 5.6 Sol 开高基本 token 开销比 Codex 少 50%,成功率也高 15%左右。下面是我跑的测试数据:[测试数据链接](https://turaai.net/benchmark)。

本来是想给我的宏命令做消融试验,但单单改执行,不改提示词,token 消耗其实只减少了 16%左右,成功率提高 11%左右。

下面是 Deep Swe 官方的解释:[官方解释链接](https://deepswe.datacurve.ai/blog/deepswe)。

简而言之,如果仅从 debug 角度来看,mini swe agent 显然比模型厂商的官方 agent harness 要好得多。

---

原文链接:[点击查看](https://www.v2ex.com/t/1232985)

评论

暂无评论。

0.035267s