The org I am at we have build an internal harness for coding, and we only hire devs that do work "AI-first". We did start hiring juniors as well as seniors; the process for either of them is the same. We drop them into any of our repos without any prior knowledge of what the project is about; and ask them to interact with the harness to build a larger feature that has been specified in project management. For the task, they can only use AI; there is no room for any manual coding at this particular org. We grade them more or less on their interaction and problem solving patterns, as well as asking us questions; and how they write prompts. The coding interview last 2 hours, and there are only 3 rounds of interviews. Before the coding interview, we do send out a brief of what it will entail.
Maybe this sounds a little strange, but this has worked exceptionally well, as we have now several juniors working with us, that do still get checked by seniors. The harness itself which holds several levels of standards, rules as well as quality gates, allows these juniors to ship out a huge amount of high quality code. We still reject about 80% of the people who interview with us, because even though they know how to code with AI, their AI interaction patterns just aren't on par with what we need. I really started to like this process. Also to clarify, the harness isn't some Coding agent setup, but rather a whole setup that works well with Codex, Claude or Cursor; its well maintained versioned and constantly improved upon, with a huge amount of automations, rules and hooks. The moment someone starts working on anything at all, this is immediately connected and tracked with project management.
That sounds like a really great idea. Would you be open to sharing more about how the tool is designed? Are there any open-source or commercial tools like that out there right now?
For some people it clearly works, for others it does not. I feel quite hopeless with a Claude subscription, but the $100 ChatGPT subscription has been a lifesaver for a lot of my work. I have very little complaints and couldn't imagine switching.
I am still not convinced there isn’t some secret basement in which each frontier lab is just orchestrating all of these agents to make their products appear much more intelligent than they are with all guard rails turned of and continuous human input.
Well let’s look at facts - provided enough compute and a goal, these system will be in a sort of loop trying out every single thing that’s in their system - they have encyclopedic knowledge and so it’s not unbelievable that a prompt which usually has a lot of implicit human rules in it can be misunderstood by AI and it just tries everything in its arsenal and we hear about the things which actually resulted in damage. I bet most of the time, they just spin in loops without achieving much if my experience with these LLMs is anything to go by. They have an important advantage in one area though, they know a lot and they can spin forget trying all sorts of combinations of things. The danger right now is probably cybersecurity, which is most likely because most orgs have historically underinvested in that area
Not the right comparison to make when comparing gold vs paper currency that can be printed at a whim. The comparison would hold if they pulled out physical fiat currency reserves.
I have yet to even try Fable or Opus 5. Just looking at the hype they put out prior to the release, then the whole way of releasing these models as well as the pricing just puts me off to get used to it and then needing it. And this while so many good models came out without any hype, botched releases and significantly cheaper. Anthropic really shot themselves in the foot. I went from using sonnet and opus models for 50-60% of my daily token usage to 10-20%
reply