AI Is Solving CTF Challenges in Minutes

(simulationslabs.com)

20 points | by therepanic 10 hours ago ago

11 comments

  • bmenrigh 10 hours ago

    I'm one of the BSidesSF CTF organizers and challenge authors (symmetric). The article is a few months old now (May) which is ancient history as far as AI advancements go. That said, I'm still emotionally coming to terms with what happened.

    Speaking from my own perspective (and not any of my challenge co-authors), I felt completely blindsided by the incredible progress frontier models made in solving CTF challenges between 2025 and 2026. We've been running the BSidesSF CTF for more than 10 years now, gaining experience on what makes a good, fun, and fair (solvable without random guessing) CTF challenge. 2026 was the first year where all of our past experience didn't seem to apply. Challenges that I designed to be hard, that I expected to take a dedicated human 10-20 hours to solve, fell to LLM automation in minutes.

    I don't know what the future of CTFs is going to be, but I wouldn't be surprised if they're largely dead in 1-2 years. A lot of the satisfaction I get from making challenges is in seeing players struggle, learn, and then eventually solve them. I'm not sure there are going to be many players willing to sink 20 human hours of their weekend into one challenge when a dozen teams using AI solved the challenge in under an hour.

    Overall I'm thrilled with the capabilities we're getting with AI, but saddened by what we're losing. I hope CTFs can somehow hold on, and that I can still get a lot of satisfaction out of building challenges and having players solve them.

    • arecsu 6 hours ago

      Maybe they will evolve to search actual exploits live in a competition or something like that :) or a marathon to research jailbreaks for certain devices and give control back to the user, like smart tvs, I don't know! Something that would be more involved and maybe real. It's just a matter of rising the difficulty, and the bright side could be that it might be used for something even more creative and good!

    • binary132 5 hours ago

      I think there will definitely still be people interested in understanding these sorts of things (that is: solving CTF’s without automatic assistance), even if at the end of the day they solve them using automation when it comes to a professional context. What I am more afraid of is people giving up and losing hope that there is any reason for humans to develop and understand systems. We need to keep in mind that even very powerful automation is still a tool for a human to make use of, and that is something I’m not confident many people are able to be conscious of. It’s a little like believing that an advanced CNC machine makes carpentry pointless to practice or understand, since now a computer can just do it.

  • simonw 10 hours ago

    This is slightly old news at this point - the BSides conference was in March, this article was published in May, we've had a whole lot of additional evidence since May that makes it unsurprising that AI can solve CTF challenges!

    Here's the repo mentioned in the story: https://github.com/verialabs/ctf-agent

    > Autonomous CTF solver that races multiple AI models in parallel. 1st place BSidesSF 2026.

    That one hasn't had any commits since March 28th. Looks like it was running Claude Opus 4.6 and GPT-5.4.

    I expect Opus 5 and GPT-5.6-Sol would be even more effective.

    (Fable 5 would refuse the challenge, Mythos 5 would undoubtedly nail it.)

  • BarryMilo 10 hours ago

    Reads like AI writing again. Can you just post the prompt?

  • elmer2 9 hours ago

    The OSCP limits tooling during the exam, even though things like Metasploit have existed for a decade.

    Companies using this for a tech interview will just need to have a proctored CTF exam, to ensure no cheating with AI.

    Bsides could split between human only and AI CTF challenges.

    These solutions aren't hard.

  • piazz 10 hours ago

    > The implications are clear: focus on the human-skills part of the job.

    The implications are not clear. They are not clear for the security people, for the SWE people, for anybody in knowledge work whose jobs are impacted.

    I wish people would stop with the “the solution is merely simply retool against the part the AI isn’t good at yet” cope and feel the enormity of the moment with humility.

    When the dust settles these jobs may not exist, or the jobs that do exist will be unrecognizable from the ones today and perhaps so qualitatively different as to no longer be attractive.

  • ComputerPerson 10 hours ago

    How are people getting models to CTF when I can't even simulate a reverse engineering attempt on an open source project?

    I'm glad I can't, but I'd like to be informed. Is it prompt engineering or just some model/harness combination?

    • elmer2 9 hours ago

      You can get access to unlocked cybersecurity models, if you are in the security industry.

      My other question is who is funding the tokens? It has to be expensive to run multple agents on complex tasks.

  • henrydark 10 hours ago
  • 12348-1632554 10 hours ago

    [dead]