Performing Excellence
“Do I need a MacBook to use Claude Code?”
Someone asked me that last week, and it isn’t the first time I’ve been asked it.
It’s a fair question. It’s also a logistics question.
Then I watched someone give a demo who had been using the same conversation thread for weeks to produce a report, loading the context of previous turns before resuming, and burning tokens the whole way. The context window was full. The evals were far from reality.
We are asking people to use AI to do the work and produce results. Increase efficiency. Reduce cycle time. Generate documents. We have selectively given tools to people to use. We have not given them any of the necessary training on how to understand and manage a context window, guidance on when to start a new conversation, or how memory persists between them.
That costs us real dollars in inefficient token usage. It also degrades the output, and people hit limits mid-task without understanding why. This tells me that they have not been taught how to think with the tool.
Then there are the people who are paralyzed. They have a license to a popular coding tool and have never opened it.
This happened to me too. I was brazenly using Lovable on my personal account to prototype ideas, and then I got a corporate license and froze. On the work account, I was waiting for inspiration, for the perfect use case that would show my best work. I was performing excellence. On my personal account, I could just build random things for a user of one.
So I wonder how much of this is really a training problem. We are measuring who is using it and who is not, and usage is the one number we can track without changing much else. We are not looking as hard at our org structures, processes, incentives, or at defining what good looks like with these tools. We have not permitted our teams to slow down on the judgment parts so they can go fast on the execution parts.
So we ask the logistics questions: which MacBook, which model, which coding tool is most efficient. We skip the judgment questions: When should a model handle a task? How would I know if the output is any good? What context is passed between one conversation and the next? And, when do I edit the markdown file myself vs. asking a model to do it?
Those are the questions I’d want us training people on. What gaps are you seeing: training or judgment?

