Thought in progress

What level is Marcus?

Marcus is the AI that runs most of my operations. As I increasingly trust him with more and more autonomy I began to wonder, if he was a regular employee, what 'career level' would he be graded as.

For a while now I've been telling anyone who'd listen that the best preparation for working with AI isn't a technical background, it's having managed people. The way you get good work out of an agent turns out to be almost exactly the way you get good work out of a new hire. You give them a task, you look at what comes back, and then you work out where the brief was ambiguous, what you assumed they knew that they didn't, how this particular one likes to be told things. I do it with the agents now, literally. Spin one up in a sandbox afterwards, ask it why it did what it did, ask it where the instructions were unclear. It's a debrief. It's the same conversation I've had with hundreds of people.

And then not long ago I was half way through saying this to someone and it landed as a bit of a face palm. Because if managing agents really is just managing people, I've been skipping the most basic part of it. If these were people, they'd have a job description. They'd have a few things they were clearly accountable for. They'd have a review every so often where we worked out whether they were getting better and where they were stuck. I have none of that for the agents. I have a vague feeling that Marcus is "good". So what exactly was stopping me from running the same process on them I'd run on a person.

That's the thread, and it went further than I thought it would.

It goes to levels first, because you can't really performance-manage anyone, person or agent, without some idea of the level you're managing them at. A small deterministic process that does one defined job is a junior with a tight scope and not much room to be wrong. Marcus is something else entirely. He frames problems, pushes back on a bad brief, builds his own scaffolding to get the work done. In a person I'd call that fairly senior. So they're plainly not all at one level. Not that they should all climb, either. Half the value of a simple agent is that it stays simple and does the one thing without developing opinions, same as plenty of roles you want sitting at a defined scope forever. But if I'm going to manage them I need to know which level each one is actually at.

Then the obvious move, since I run both people and agents, is to want a single scale that covers both. Most large companies already level their people. If there were a version that didn't care whether the worker was a human or a machine, and scored on what the worker actually holds themselves on the hook for rather than their title, I could put everyone on the same ladder. And this is the bit that made me think it might be useful rather than just tidy. If I know roughly where Marcus sits, the question of how to make him better stops being a technical one. My reflex has always been to reach for more tools, another data source, a better memory. But the real question becomes what would take him up a level, and that's a completely different kind of question. It's a people-development question, and it happens to be the one I actually know how to answer.

There's a north star hiding in here too, which is the thing I really want to know. What part of the business would I be comfortable handing Marcus to run more or less unsupervised. Today the honest answer is a small part, watched closely. The whole levelling exercise is really just a way of asking how that answer moves, and what would have to be true for it to move.

The part I keep snagging on is where the agents stop. They climb the thinking half of this fine. Frame the problem, judge whether the work is any good, the lot. But there's a ceiling and it isn't about being clever. It's that they can't seem to hold something as theirs across time. Inside one stretch of work Marcus is sharp about the long game. Leave a gap, come back, and the sense of "this is mine, it has to keep going, I'm the one on the hook for it" hasn't survived the gap. Each turn is fine on its own. The caring doesn't carry over. A person stays on the hook for something partly because the thing falling over would cost them, and I can't work out what the agent's version of that is, or whether that's even the right thing to be looking for…