I truly empathize with this idea. It does feel like things are improving. I was scanning the code generated and felt the same way. Like, woah, maybe I'm in the clear and I can hand the reigns over. Do I even need to look at the code anymore?
Then I ran into a few bugs and, as I dug in, what initially looked like reasonable code suddenly seemed strange when looking closer. After some help from the agent to understand the intent, it became clear the design was poor, explained a few failures, and resulted in an order of magnitude more (reasonable looking) code to compensate which now I have to sift through.
To make matters worse, I could not get the agent to divorce itself from its wrong decisions, even with my explicit instruction detailing how the code should read. It would keep rewriting the same bad ideas then go into other distracting tangents which I'd have to correct. It felt just like how conversations in ChatGPT that would get stuck once it was in the context making me imagine we're still driving the same car but with a new paint job.
I finally wrote the code myself. It was more helpful after that. Hopefully, I'm able to resist the temptation and stay vigilant in my reviews. For context, I am using Codex Astra XHigh but maybe Opus 5.5 really is better?
Astra has been quite good as well but I prefer Opus 5.5.
Definitely worth spending more time on the plans & auditing for quality. Once the plan is ready, I find that adherence to plan is quite solid.
Another angle I look at it is that the code is going to be imperfect regardless of review. The main criteria is whether the code is producing the business outcome that you are looking for. Does it fulfil the user stories? Is it performant? Do you have enough pre-release checks and balances to make sure its not going to cause issues.
Maybe it depends on what you are working on. I work on platforms and engineering tooling. I cannot say the agent can be trusted yet in my experience. And I’d say I’m pretty advanced agent user compared to my peers. The majority of the work I do I have to steer and correct a lot.
For example, recently I’ve been adding skills for our team. I added evals. If you ask an agent to make evals it’ll always try to bias and overfit just to make tests pass. No matter if I have it stated in every possible place not to do it and Codex instructed to catch these cases. It doesn’t work. Lots of sloppy evals get written and skills get extremely specific instructions to pass specific tests.
I think if you are implementing trivial features in a green field product it can work to some degree for some time. But eventually it’s going to deteriorate into a mess. Death by 1000 paper cuts.
Will be interesting to hear an update after two months. I think two weeks is too little to conclude that it works.
My experience and what I’m concerned about is that eventually more time is spent refactoring to enable new features than writing the features themselves. In my experience this becomes worse with time as the models are reluctant to remove behavior, so the code base will grow to support some version of the old behavior together with the new behavior.
This was exactly me a few weeks ago. I felt unbeatable, things that I planned for several weeks finished in a week. I asked the AI about everything I cared about, special cases, how the code worked together, etc... 8 merge requests for my team to review. Man, it felt good.
Hats off to my team, they did review the code. And it was embarrassing. I could not answer any questions asked by my team without looking at the diff. It was supposed to be "my changes", as we have agreed to the team. You can do everything with AI, but you own the changes. I did not. Why do we need to remove duplicates and call GraphQL in batches of 25 IDs when it's for 3 items displayed on the page? Did I have that much distrust in our Ops team to not believe that they can choose 3 unique IDs without having duplicates and know the difference between 3 and 25?
Use AI for anything you want but make sure you own the changes. Otherwise, you're just a meat proxy. Don't be a meat proxy like me. Be better. This world needs you more than ever to own it.
This is an LLM-written marketing blog-post about vibe-coding and token-maxing.
I read big claims, no evidence.
"AI code is non-deterministic and has risk. But human code is non-deterministic too, and it has the exact same risk."
No, that's not true. LLM-written code has very different risks, like completely misunderstanding the requirements, adding in hallucinated features, and losing sight of what the codebase actually does (massive tech debt).
I should ask, who writes the unit tests? And how do you write the unit tests beforehand? What if you need to prototype in order to figure out the shape of the API? The setup used here is so far removed from any software safety or software quality concerns, it would be hilarious if it weren't sad.
Stop posting marketing content void of any real information, please.
I truly empathize with this idea. It does feel like things are improving. I was scanning the code generated and felt the same way. Like, woah, maybe I'm in the clear and I can hand the reigns over. Do I even need to look at the code anymore?
Then I ran into a few bugs and, as I dug in, what initially looked like reasonable code suddenly seemed strange when looking closer. After some help from the agent to understand the intent, it became clear the design was poor, explained a few failures, and resulted in an order of magnitude more (reasonable looking) code to compensate which now I have to sift through.
To make matters worse, I could not get the agent to divorce itself from its wrong decisions, even with my explicit instruction detailing how the code should read. It would keep rewriting the same bad ideas then go into other distracting tangents which I'd have to correct. It felt just like how conversations in ChatGPT that would get stuck once it was in the context making me imagine we're still driving the same car but with a new paint job.
I finally wrote the code myself. It was more helpful after that. Hopefully, I'm able to resist the temptation and stay vigilant in my reviews. For context, I am using Codex Astra XHigh but maybe Opus 5.5 really is better?
Astra has been quite good as well but I prefer Opus 5.5.
Definitely worth spending more time on the plans & auditing for quality. Once the plan is ready, I find that adherence to plan is quite solid.
Another angle I look at it is that the code is going to be imperfect regardless of review. The main criteria is whether the code is producing the business outcome that you are looking for. Does it fulfil the user stories? Is it performant? Do you have enough pre-release checks and balances to make sure its not going to cause issues.
Amazingly. This sounds like a job I can hire people to do really cheaply.
Maybe it depends on what you are working on. I work on platforms and engineering tooling. I cannot say the agent can be trusted yet in my experience. And I’d say I’m pretty advanced agent user compared to my peers. The majority of the work I do I have to steer and correct a lot.
For example, recently I’ve been adding skills for our team. I added evals. If you ask an agent to make evals it’ll always try to bias and overfit just to make tests pass. No matter if I have it stated in every possible place not to do it and Codex instructed to catch these cases. It doesn’t work. Lots of sloppy evals get written and skills get extremely specific instructions to pass specific tests.
I think if you are implementing trivial features in a green field product it can work to some degree for some time. But eventually it’s going to deteriorate into a mess. Death by 1000 paper cuts.
Will be interesting to hear an update after two months. I think two weeks is too little to conclude that it works.
My experience and what I’m concerned about is that eventually more time is spent refactoring to enable new features than writing the features themselves. In my experience this becomes worse with time as the models are reluctant to remove behavior, so the code base will grow to support some version of the old behavior together with the new behavior.
This was exactly me a few weeks ago. I felt unbeatable, things that I planned for several weeks finished in a week. I asked the AI about everything I cared about, special cases, how the code worked together, etc... 8 merge requests for my team to review. Man, it felt good.
Hats off to my team, they did review the code. And it was embarrassing. I could not answer any questions asked by my team without looking at the diff. It was supposed to be "my changes", as we have agreed to the team. You can do everything with AI, but you own the changes. I did not. Why do we need to remove duplicates and call GraphQL in batches of 25 IDs when it's for 3 items displayed on the page? Did I have that much distrust in our Ops team to not believe that they can choose 3 unique IDs without having duplicates and know the difference between 3 and 25?
Use AI for anything you want but make sure you own the changes. Otherwise, you're just a meat proxy. Don't be a meat proxy like me. Be better. This world needs you more than ever to own it.
I tried letting my agent take full control. You won't believe what happened next.
The subtitle was so Claudish I immediately bailed on reading this.
This is an LLM-written marketing blog-post about vibe-coding and token-maxing.
I read big claims, no evidence.
"AI code is non-deterministic and has risk. But human code is non-deterministic too, and it has the exact same risk."
No, that's not true. LLM-written code has very different risks, like completely misunderstanding the requirements, adding in hallucinated features, and losing sight of what the codebase actually does (massive tech debt).
I should ask, who writes the unit tests? And how do you write the unit tests beforehand? What if you need to prototype in order to figure out the shape of the API? The setup used here is so far removed from any software safety or software quality concerns, it would be hilarious if it weren't sad.
Stop posting marketing content void of any real information, please.
[flagged]
[flagged]
[flagged]