Skip to content

Deep dive

From GPT-2 to AI agents: when answers stopped being enough

A personal reflection on GPT-2, chatbots, automation, and AI agents, through one broken app and the question of how much work we can hand over.

24 min read
  • AI
  • AI agents
  • AI infrastructure
An unavailable app, a chat assistant suggesting checks, and a terminal connected by a repeated handoff loop.

Here’s the scene I keep coming back to.

A small app has stopped working after a deployment. Nothing dramatic. A page that won’t load, a few error messages, and a change that was supposed to be harmless.

I open a chat window, paste the error, and ask what’s wrong.

The answer sounds sensible. Check the configuration. Look at the database connection. Compare the latest deployment with the previous one.

Useful. But now I have to go and do all of that.

Open the terminal. Find the logs. Copy something back. Explain what I found. Try another suggestion. Paste the next error.

Somewhere in that back-and-forth, the question changes. I’m no longer wondering whether AI can explain the problem. I’m wondering why I’m still carrying every piece of information between the AI and the problem.

Why am I the copy-and-paste layer?

That imaginary broken app is going to stay with us through this article. It’s a simple way to think about a much bigger journey: from GPT-2 continuing a piece of text to systems that might one day look after a piece of work without needing me to push every step forward.

And it’s how I ended up asking what comes after agentic AI.

Illustration 01: The copy-and-paste layer

1. Before we asked it to fix anything, we asked it to keep writing

Looking back at GPT-2 in 2019, I have to reset my expectations a little.

It’s easy to judge an older model against the things we now want from AI. Could it investigate a repository? Could it use a browser? Could it recover from a failed deployment?

But that wasn’t the interesting question at the time.

GPT-2 was a language model trained to predict what came next in text. Give it a starting point, and it could produce an extended continuation. The striking part was how much recognisable structure could emerge from that setup: paragraphs, explanations, dialogue, passages that seemed to know what kind of writing they belonged to.

I’m looking back here, not claiming I was running every early model as it arrived. What interests me is the distance between that starting point and the expectations we now place on the same basic interaction: type something into a box and wait.

With GPT-2, the question was closer to, “Can this system continue something that makes sense?”

With an agent, it becomes, “Can this system work out what to do, do it, and check the result?”

Those are very different jobs.

Illustration 02: More text, same app

There is something from the GPT-2 stage that I don’t want to lose sight of, though.

A model can write the words “the problem is fixed” without fixing anything. It can describe a database connection without opening one. It can produce the shape of a convincing explanation without having seen the relevant evidence.

That gap doesn’t disappear just because the prose gets better.

The rest of this story is really about trying to close it.

2. GPT-3 made me think about the prompt differently

GPT-3 arrived in 2020, and the prompt became a more interesting place to work.

You could describe a task, provide a few examples, and ask the model to continue in the same pattern. The same underlying model could be used for quite different kinds of work, depending on what you put in front of it.

I find that shift easy to underestimate now. Instead of building a separate interface for every small language task, a developer could experiment by changing instructions and examples.

Summarise this. Classify that. Rewrite this paragraph. Follow this pattern.

It wasn’t the model permanently learning a new skill every time it saw an example. The examples were part of the current context. Still, the flexibility was useful.

Illustration 03: One model, different prompts

Then code generation made the whole thing feel closer to the work itself.

The early Codex model in 2021, which powered the early GitHub Copilot experience, helped connect ordinary language with code. That matters because code can leave the page. It can be run.

A paragraph telling me how to process a file is one thing. A function that expresses the transformation is another.

But put that function back inside our imaginary app and the missing pieces become obvious. Which file should it live in? Does it match the existing code? Are the dependencies available? Does it actually solve the error? What else might it break?

The model could help produce a piece of the answer. I still had to turn that piece into working software.

That’s the first thread I notice running through this history: the output gets more useful, so I immediately start noticing the work around it.

3. ChatGPT made the conversation feel natural. The work was still elsewhere

By 2022, instruction-following work such as InstructGPT was pushing models towards a more useful relationship with a request. Rather than just continue a piece of writing, try to respond to what the person is asking for.

Then ChatGPT arrived in November 2022 and put that interaction inside a conversation.

For me, the important part of chat is the permission to be unfinished.

I don’t need to get the entire request right in one attempt. I can ask something, see the answer, realise what I meant, and adjust. “No, make it simpler.” “Use Python.” “Explain that part again.” The conversation becomes a way of finding the question as well as answering it.

That is genuinely useful. I wouldn’t dismiss it just because agents can do more kinds of work.

But let’s return to the broken app.

3.1 The assistant can explain. I still have to investigate

In a chat-only setup, I provide an error about a failed database connection. The assistant gives me possible explanations.

Perhaps the database is unavailable. Perhaps the address is wrong. Perhaps a setting changed during the deployment.

Then the conversation stops and I start moving.

I fetch another log. I check the configuration. I decide which suggestion is worth trying. If the assistant needs the previous deployment’s settings, I go and find them.

There is intelligence helping with the investigation, but the investigation still lives largely in my head and my hands.

Illustration 04: The chat investigation loop

3.2 The chat window isn’t the capability

This is where my original ladder starts getting a little messy.

I might say “chat interface, then chatbot, then automation” as though each one replaces the last. But the interface is just where I speak to the system. It doesn’t tell me what the system can do.

Two products can have almost identical chat boxes. One can only answer. The other can inspect files, run tools, and check results. And an agent doesn’t need a chat box at all; an event could start its work.

So when I think about this progression now, I try to look behind the conversation.

Who is deciding the next step? Who is carrying it out? Who notices when it didn’t work?

Those questions tell me much more than the shape of the input box.

Illustration 05: Same chat, different capability

4. The first thing I’d automate is the boring handoff

Once I notice how much of the job is moving information around, automation becomes an obvious place to start.

For our app, I could define a workflow that runs when an alert arrives. It collects the relevant logs, captures deployment details, removes the noise, asks a model for a summary, and opens a ticket.

Now I don’t have to collect the same evidence by hand every time.

There is AI in the workflow, but the workflow itself is mostly something I’ve already decided. I know which information to gather. I know where the summary should go. I can specify what happens when a step fails.

This is what I mean here by AI automation.

Illustration 06: Automate the handoff

I actually like this stage more than the usual “automation is basic, agents are advanced” story allows.

A predictable workflow can be exactly what I need. If the task is known and repeatable, I don’t automatically want a model inventing a new route through it. I want the task done without making it my problem again.

But there is a limit.

Suppose the summary says the database connection failed, and the prepared workflow has collected everything it was told to collect. It hasn’t compared the new configuration with the previous deployment, because I never added that step.

The workflow has finished. The investigation hasn’t.

I could keep adding branches. Sometimes that is the right answer. But eventually I might want the system to notice what information is missing and choose a useful next check itself.

That’s the point where an agent starts making sense to me.

Not because the name is more impressive. Because the next step depends on what we find.

5. The interesting moment is when the next step isn’t mine

Function calling made the connection between models and software more explicit in the 2023 API landscape. Instead of only returning ordinary text, a model could return structured arguments for a function the developer had made available.

The application still had to validate and execute the request. The model didn’t magically gain access to every system it mentioned.

That separation is easy to miss, but I think it’s the hinge of the whole story. A model can ask for something to happen. The surrounding software decides whether the request is valid and allowed, then reports what happened.

Illustration 07: Proposal is not execution

5.1 Let the investigation change direction

Back at the app, imagine I give an agent read-only access to the logs and deployment configuration.

It sees the connection error. Its first hypothesis is that the database is unavailable. It checks the evidence it is permitted to inspect. That explanation doesn’t fit well enough.

So it looks at what changed in the release.

In this imagined case, the application is pointing at the wrong database address. The database itself wasn’t the problem. The new configuration was.

What matters here isn’t that the agent guessed correctly on the first attempt. It didn’t. What matters is that it could collect another piece of evidence and change its mind without waiting for me to write the next prompt.

The loop has moved.

Previously, I read the response, chose an action, ran a tool, and returned with the result. Now the system can take on some of that cycle within the access I’ve given it.

Illustration 08: Follow the evidence

5.2 Finding the cause isn’t the end of the story

An agent finding the likely cause feels like a big step forward. But I don’t want the story to end with a confident explanation.

It could prepare a proposed configuration change. It could check the change in an isolated environment. It could show me what it intends to do and what evidence supports it.

Whether it should touch production is a separate decision.

And even after an approved change, the app needs to be checked. A successful tool call tells me that a command completed. It doesn’t necessarily tell me that a user can now use the service.

This is the part that makes agents interesting to me: they bring the question of completion into the system.

Not just “What should I say next?” but “What would convince us this actually worked?”

Illustration 09: A fix needs verification

6. Then I got stuck on the words “agentic AI”

Once an agent can investigate, plan, use tools, and change direction, what exactly am I adding when I call the system agentic?

I don’t think there is a clean new species hiding behind the phrase.

The distinction that works for me is simple: an agent is the thing doing the work; agentic describes the way the work is being handled. Goal-directed, able to choose actions, responsive to feedback.

A system can have one agent or several. It can mix model decisions with fixed workflows. It can give the model freedom over an investigation while keeping deployment under a very ordinary approval process.

That mixture feels more realistic to me than an architecture where every box has to be an agent.

Illustration 10: Different kinds of control

I started with a neat sequence because neat sequences help me think. Chatbot, automation, agent, agentic AI, then whatever comes next.

But the closer I look, the more useful it becomes to pull that sequence apart.

A chat interface isn’t a level of autonomy. Several agents aren’t automatically more capable than one. A long-running process isn’t necessarily learning. A system can become much better at a narrow job without becoming generally intelligent.

So I still use the progression as a story about increasing delegation. I just don’t mistake it for a fixed technical ladder.

And that changes the question from “What is the next label?” to something I can actually build around:

What part of the work am I still having to hold together?

7. The model gets the attention. The surrounding software gets the difficult jobs

This is where the history from 2024 into 2026 becomes especially interesting to me.

MCP appeared as a way to standardise connections to tools and context. A2A addressed communication across agent boundaries. Coding-agent products began packaging repositories, tools, and execution environments into the experience rather than leaving all of that outside the model interaction.

The 2025 cloud coding agent called Codex was part of that shift. Same familiar name as the earlier coding model, but a different kind of product experience: work happening inside an environment, not just code appearing as an answer.

I don’t read those developments as the model suddenly becoming the entire computer.

I read them as the surrounding software catching up with what we’re asking the model to do.

Illustration 11: Inside a coding-agent workspace

Think about what happens if the investigation of our app lasts longer than one session.

The system needs to remember which explanation was ruled out. It needs the exact change it proposed, not a vague memory that it “looked at configuration”. It needs to know whether approval is pending. If something crashes, it needs to pick up without repeating an action that already happened.

None of that gets solved by writing a more enthusiastic system prompt.

An agent harness is the software that makes this possible: calling the model, providing context, dispatching tools, and keeping the process moving. Work on longer-running agents has explored progress files, Git history, handoffs, and durable session records for precisely these kinds of problems.

The details keep changing as models improve. I don’t assume every piece of scaffolding will be needed forever. But the questions remain very practical.

What happened? What is still unfinished? What can safely happen next?

This is also where my interest in platform engineering starts to overlap with my interest in AI. There is a lot of familiar work hiding under the new vocabulary: state, isolation, retries, permissions, observability.

The model has become part of a system that can fail in more ways than simply giving a bad answer.

Illustration 12: Resume from a durable record

8. What happens when I close the laptop?

This is the next step I find myself thinking about.

So far, I’ve asked the system to investigate a particular problem. It does some work, returns a result, and the interaction ends.

But what if the job doesn’t end with the conversation?

For our imaginary app, I might want a system to keep track of the incident, wait for my approval, carry out an authorised change, check the outcome, and bring the unresolved parts back to me.

Not because it has permission to do anything. Because it has a specific piece of work to carry across time.

8.1 I don’t need it thinking every second

The phrase “always-on agent” makes it easy to picture a model endlessly generating thoughts in the background.

That isn’t the version I want to build.

Much of the time, the right action would be to wait. Wait for a deployment to finish. Wait for an approval. Wait for enough observations to tell whether the change helped.

The task can stay alive without the model constantly running. A durable process can hold the state and wake the model when there is a decision worth making.

Illustration 13: Waiting is part of the work

8.2 The promise needs boundaries

“Look after my app” sounds appealing until I try to define it.

Can the system restart something? Change a setting? Spend money? Contact a customer? Touch production? What should it do when the evidence is unclear?

Those questions aren’t annoying details that get in the way of autonomy. They’re what make the delegation meaningful.

I can imagine saying: investigate alerts, prepare fixes in a separate environment, and ask before making production changes. Stop when the investigation budget is exhausted. Keep a record that I can inspect.

That’s a much more useful agreement than “be autonomous”.

The shift I’m imagining is from giving the system an isolated task to giving it a limited responsibility. It can keep that responsibility until the work is completed, cancelled, or handed back.

That doesn’t have to mean AGI. It might simply mean a much better way of keeping a small, well-defined job from falling between conversations.

Illustration 14: A bounded responsibility

9. Maybe I need several agents. Maybe I just need a better tool

The next temptation is obvious.

If one agent can help, why not a team? One for investigation, one for code, one for tests, another to review the others.

I understand the appeal. It’s easy to look at a software project and start mapping human job titles onto model calls.

But when I put that team around our broken app, I also see the possible mess.

The investigation agent changes its hypothesis. The coding agent is still working from the old one. The testing agent checks the wrong behaviour. A supervisor spends another round explaining what changed.

Now the system has work about the work.

Illustration 15: Coordination has a cost

There are jobs where separate workers make sense. One could inspect deployment changes while another checks a separate source of operational evidence. Independent research questions can be explored in parallel and combined later.

But two agents editing the same tightly connected code might spend more time reconciling their work than they save.

And sometimes the supposed specialist doesn’t need to be an agent at all. To find a symbol, use a search tool. To check whether a build passes, run the build. I don’t need a second model to narrate either operation.

So I don’t see multi-agent systems as the mandatory thing after agentic AI. I see them as one way to divide a job when the division actually helps.

The question I’d ask before adding another agent is embarrassingly ordinary: what work will this one own, and how will the others know whether it did that work properly?

10. “Next time, don’t make me explain this again”

Once a system can carry a task across time, I immediately want something else from it.

I want the next attempt to benefit from the last one.

For our app, that might mean remembering the right build command, where configuration lives, or why a previous approach was rejected. I don’t want every new session to rediscover the same basics.

That sounds like learning. But I think the word hides several different things.

10.1 Sometimes it just needs good notes

A stored project fact can be retrieved later. That changes the information available to the system, without necessarily changing the model itself.

There is nothing disappointing about that. I use notes for the same reason: I don’t want to reconstruct everything from scratch.

What bothers me is the idea of a system remembering something without remembering why it believed it.

“The database is always the problem” would be terrible memory. “In this deployment, this configuration change caused this connection failure” is a much more useful record.

And even that needs a scope and a date. Our app can change. Yesterday’s fix can become tomorrow’s misleading clue.

Illustration 16: Keep useful project notes

10.2 Sometimes the method itself should improve

Suppose the system’s first approach is to read half the repository before finding the relevant configuration. A later approach starts with a targeted search and reaches the useful evidence with less wandering.

I’d want a way to compare those approaches.

Not just on this one incident, though. A shortcut that happens to work here might skip important evidence somewhere else. I want to check it against other tasks before turning it into the new default.

This is the kind of self-improvement I find concrete enough to reason about: propose a change to the method, evaluate it, and keep it only when the evidence supports doing so.

Illustration 17: Evaluate a better method

Systems such as AlphaEvolve make this idea easier to picture. Language models propose candidate programs, evaluators test them, and the search builds on useful candidates. The ability to evaluate the result is central.

I find that more interesting than the vague promise that an AI will simply “make itself smarter”.

Changing retrieved notes, changing a workflow, and updating model weights are different operations. None should be confused with proof of unlimited self-improvement.

The version I’d trust has something outside the candidate saying whether the change helped. It doesn’t get to mark its own homework by quietly making the test easier.

Illustration 18: Three different kinds of change

11. Then the boundary might move beyond my own system

Once I imagine agents carrying defined responsibilities, it becomes natural to wonder how they might work with other systems.

Not another sub-agent that I control. A separate service with its own tools, rules, and owner.

Perhaps my system needs a specialist analysis. Perhaps it needs a dataset. It could describe the required result, request the work, and check what comes back.

That’s the version of an “agent economy” I can start to picture. Less a room full of digital people negotiating, more software arranging clearly specified services on someone’s behalf.

Illustration 19: A request across service boundaries

But this is also where the questions get harder.

Who authorised the request? What counts as delivery? What happens if the result is wrong? Can the service see private data? Can it commit money? Who resolves a disagreement?

A communication protocol doesn’t answer all of that. It helps systems talk; it doesn’t settle what they should trust.

I can imagine a business eventually built around more of these arrangements: systems handling routine operational work, people setting direction, reviewing exceptions, and deciding which tradeoffs are acceptable.

I don’t take that to mean a business with no people, or no responsibility left for them. If anything, I think the responsibility needs to become more visible as execution becomes less manual.

Someone still has to decide what is worth doing. Someone has to care when the system produces the wrong outcome for a customer, even if every internal step says it succeeded.

Illustration 20: People keep responsibility visible

And I would keep all of this separate from a prediction about general intelligence. A system could become excellent at coordinating a narrow set of services and still struggle badly outside them. Physical action brings another set of demands again.

I see branching possibilities here, not a staircase with AGI waiting neatly at the top. I don’t know when, or whether, every branch becomes practical.

12. The part I keep coming back to is trust

There is a point in every version of this story where the system says it’s finished.

For a chatbot, I can read the answer and decide whether it helped. For a system that changes things, “finished” becomes a much more demanding claim.

Return to the app one last time before we talk about the future.

The configuration has been changed. A command returned successfully. The agent writes a tidy summary.

Would I close the incident?

Not yet. I would want to see the application working. I would want to know what was checked, what wasn’t, and whether the change created another problem.

That doesn’t mean I need impossible certainty. It means the checks should match the thing I asked the system to achieve.

Illustration 21: Completion needs evidence

12.1 What happens when the boring things go wrong?

Suppose a tool times out after the system requests a change.

Did the change fail? Or did it happen, with the response getting lost on the way back?

Blindly trying again could repeat the action. Starting the entire investigation over could waste everything already learned. Pretending it succeeded would be worse.

The system needs enough recorded state and access to evidence to work out what happened, or enough restraint to ask for help.

That’s the kind of reliability I care about. Not a demonstration where nothing goes wrong, but a process that doesn’t become nonsense when something does.

Illustration 22: Check before retrying

12.2 I want to count the work that comes back to me

The cost question follows directly from that.

A cheap model call isn’t a cheap solution if I spend the next hour repairing what it did. A fast answer isn’t a fast task if it sends the investigation in the wrong direction.

I’d want to count the whole job: model calls, tools, retries, infrastructure, review, and cleanup. Then I’d look at what was actually completed to an acceptable standard.

And I would pay attention to how much supervision the system still demands. If I have to watch every step, explain every transition, and independently reconstruct every result, I may have automated less than the demo suggests.

That is the test I’d use for the next stage of these systems. How much useful responsibility can they carry without quietly handing the difficult parts back to me?

13. Strangely, the future I want might use less AI in some places

This is where my thinking keeps landing.

The more capable the overall system becomes, the less reason I see to make the largest model do every little operation inside it.

To find a function, use search. To extract a field, use a parser where the structure allows it. To check a build, run it. To format code, use the formatter.

Let the model deal with the parts that genuinely need interpretation: the ambiguous requirement, the unexpected evidence, the choice between plausible approaches.

I don’t need it to perform being intelligent while ordinary software waits beside it, unused.

Illustration 23: Use the right tool for each step

Smaller models might help with narrow decisions. Multiple agents might help with independent work. A fixed workflow might handle most of a task before a model ever needs to become involved.

But I’d want to measure those choices rather than assume every extra layer saves time. Routing can fail. Handoffs can lose information. A clever architecture can become an expensive way to do something simple.

This is the infrastructure work that interests me: giving the system the right context, making tools efficient, keeping execution recoverable, and knowing when a model call isn’t needed at all.

It’s less spectacular than imagining an autonomous digital company. It is also much closer to something I can break down, build, and test.

I want to spend my own attention where it changes the outcome. I want the system to do the same with its expensive reasoning.

14. Back to that broken app

At the beginning, I was moving an error message into a chat window and carrying suggestions back into a terminal.

The app wasn’t working, and I was holding the entire investigation together.

Now imagine a more useful arrangement.

The system gathers the evidence it is allowed to access. It investigates. It records what it found. It prepares a change and asks where it needs permission. After an authorised action, it checks whether the app actually works. If it can’t finish, it tells me what is unresolved without pretending otherwise.

There are still decisions for me. There should be.

But I am no longer required to be the connection between every pair of steps.

Illustration 24: One meaningful decision returns to you

That’s how I think about the journey from GPT-2 to whatever comes next.

First, I wanted the text to make sense. Then I wanted the answer to be useful. Then I wanted the system to help with the work around the answer.

Now I want to know how much of that work I can hand over without giving up visibility or judgement.

Maybe the next phase gets another name. Autonomous systems. Adaptive agents. AI-native operations. I can see useful ideas inside all of them.

But the moment I’m looking forward to is smaller and more practical than a new label.

It’s opening the laptop, seeing what happened, understanding what still needs me, and not having to restart the whole story.