The more I use Opus 5.5, the more I realize how far behind OpenAI models are.
Not just coding. Everything.
I sent GPT-6 Astra my V02 max lab test results, and it gave me a headache. 200 sentences of: "your report shows this, I think that, therefore it's inconclusive".
On the
I don't think writing a spec and splitting it into tickets is how you should use agents since Sol 5.6/Opus 5. Just tell the agent what you want to do and it will do it.
If the task is large, discuss with the agent how to split it in several shippable PRs and do it one at a time.
mattpocock/skills v1.3 is out!
- /pr (new) writes easy-to-read PR bodies, showing hard evidence that the changes work and assessing merge risk
- /implement-spec (new) takes in a spec and tickets, and implements them with subagents
- CONTEXT.md renamed to GLOSSARY.md
- /retro
i am not convinced that Astra is really that capable of a model - at least not really any different than Sol
like once you give _any_ of these models a task that is very complex, you find them all flailing around, often in pretty similar ways
e.g. Astra likes to go off the
I still think that best code is no code and constantly amazed by how much cruft agents add. The most usual case is agents overthinking a solution by not seeing where the real problem is (hint, it's always somewhere upstream) due to their local optima RL induced thinking.
I thought that most of us were driven by the feeling of finding the 20 lines that replace 200 convoluted lines, finally directly representing the solution instead of gesturing at it vaguely through a cloud of indirection. But the zeitgeist suggests it's weird to care about that?