AIAI EngineerJul 2, 2026· 20:13

The Prompt Is Still a Punch Card - Ted Johnson, JoinIn AI

Ted Johnson argues that AI interfaces still use the batch protocol of punch cards, forcing humans to adapt to machines rather than the reverse. He introduces three concepts—channel, expression, protocol—to show how expression exploded with LLMs but the protocol remained static. Examples include voice mode misinterpreting a side conversation, PersonaPlex's turn-taking, and a meeting AI that follows group dynamics without prompts. Johnson calls for designing interfaces where AI understands timing, ambiguity, and shared context, moving from prompting to genuine communication.

Transcript

Intro0:00

Ted Johnson0:07

I'm sure that sometime in the last few hours most of you did this: you typed a request into a small box to a superintelligence, and then you waited. You watched the cursor blink, maybe a little throbber, cycled through clever gerunds like "hullabalooing," "tomfoolering," and "philosophizing" to hide the weight.

Maybe it gave you what you wanted, maybe you rephrased it and tried again. It all felt completely normal. I want to spend the next 20 minutes making prompting feel unfamiliar and strange again. I'm Ted Johnson, co-founder of JoinIn AI.

During my 25-year career building enterprise software, collaboration systems, and AI-enabled interfaces, I've always focused on human interaction. I've also been following AI for two decades, including back to the far less impressive GPT-1 and 2. And when ChatGPT arrived, I felt two things at once: first, an unsurprisingly amazement, knowing the world would never be the same.

Followed by, actually, surprising disappointment I couldn't shake. This disappointment turned into an observation that started a company, JoinIn AI, and that I keep coming back to, which is: why do we still have to "learn" AI? Why does something this powerful so often feel unnatural to use?

Three Concepts1:29

Ted Johnson1:29

Here's the path we'll take to answer that. We'll start with the most familiar computer interface and make it strange again. Then I'll give you three key concepts: the channel, the physical transport that carries your intent; expression, the range and richness of meaning the channel can carry; and the protocol, the shape or rules of this exchange.

I'll use those three concepts to show you that the prompt is our present-day punch card. We'll share examples of the ways interfaces could progress, and we'll wrap with some practical advice for AI and human-centered design. Everyone knows what this is: the keyboard.

It's everywhere, and it feels completely normal or natural. But it isn't. We all had to take lessons. We all had to practice. And I say this as someone who loves keyboards, but it seems that we've been trying to fix them as long as they've been around.

People have tried more efficient layouts, like Dvořák or Colemak, to save their fingers some work. A more extreme example: some keyboard enthusiasts refused to squander their two digits on the spacebar, giving them 4 or 8 or 10 keys to press with that efficient thumb of theirs.

And what do we even mean by "the keyboard"? Here's the patent drawing for the layout we use every day. This patent's from about 1860. And my personal favorite: it just as could have easily been the Hansen writing ball, which looks anything but like a way you want to talk to a superintelligence.

We carry these legacies of an arbitrary input device, designed under constraints that haven't existed for a century, and we put it between ourselves and the most capable machines ever built. Nobody alive chose it. We all inherited it, and then we stopped noticing.

Channel3:18

Ted Johnson3:18

So that's the first idea: the channel, the medium an interface gives you to work in. A keyboard is a channel. A microphone, channel. A screen, a punch card, a prompt box, all channels. And channels matter because each one can physically carry a different kind of signal.

For example, text is a stream of discrete symbols. Voice can carry timing, pitch, hesitation, and words. A diagram can carry spatial relationship all at once. But these are differences in what the medium can transmit: its bandwidth, not the differences in the meaning.

And carrying more signal isn't the same as the machine understanding any of it. That's a separate question. That's the next idea. Humans use all these channels constantly, without thinking. We never pick one channel and force everything through it.

That would be absurd. And yet, that's what we ask people to do with machines, over and over. Hold on to the word "channel," because here's the plot twist. With AI, the channel never really changed. You're still typing into a box, but what you were allowed to push through it was about to.

For the first time, the computer channels carry rich, complete human language. Notice I said "what it carries," not "what it is." You're still typing into a box. You're still hitting the Submit button. The keyboard didn't change. What improved is the range of what you're permitted to express through it.

There's the second idea: expression. How much of what you actually communicate or mean will go through the interface? That's the second idea: expression. How much of what you actually communicate or mean will the interface let through? Here's an example of expression progress over time with computers.

Expression4:44

Ted Johnson5:04

Starting with assembly, which gave you an instruction set, a few dozen op codes. Then came the commands with the shell, inputs, flags. Then modern programming languages gave you primitives that you could compose. While powerful, step by step, each one of these is a fixed vocabulary, requiring you to express your intent by choosing from a menu the machine will accept.

Natural language blew that menu open. For the first time, you can say almost anything the way you'd say it to another person. And on the expression-one axis, the leap is real and enormous. There's an ocean of meaning in an ordinary human request: context, nuance, intent, all things we've never had to spell out to each other.

For the first time, you can say almost anything the way you'd say it to another person. For the first time, a machine can take it in. So here's what should bother us as engineers and designers, as it's bothered and inspired me: the channels for computers have been the same for 15 years, some 180 years.

Now, with AI, we've poured an ocean of expression into it. So why does it feel like we're still sipping through a straw and struggling to learn how the AI thinks? Because there's a third idea underpinning the other two, and it's really the one that hasn't kept up: the protocol, the rules you follow, and the shape of the interaction itself.

Protocol6:26

Ted Johnson6:26

Channels stayed the same. Expression exploded in the last three years with LLMs. But the protocol, prompting, is the protocol of a punch card. And the punch card's protocol is good old batch. Here's what punch card batch meant: you sat down, away from the machine, carefully encoded your entire request in advance, carried your deck to the operator, you submitted the job, and then you waited, sometimes hours, sometimes overnight.

Then you read the printout, found one thing that was wrong, fixed it, resubmitted it, and waited again. The machine never engaged with you while you were thinking. It engaged with the finished package after the fact. Now let's look at the prompt.

Assemble the whole request, submit it, wait, read what comes back. Something's off, assemble it again, submit again, and wait. We have to acknowledge that there are features, interactive features, improving this. You can ask for updates. You can ask for summaries of what was done.

But at the end, it's still batch with interactive sprinkles. It's the same protocol. We shrank the wait time from overnight to a few seconds, or a few minutes, and the speed fooled us into thinking that it had become interactive.

It hasn't. It's still batch. You still package a complete turn before the machine is allowed to participate. We learn tricks, send tips to use code skills, or rewrite prompts a certain way to manage this. And speaking doesn't change it.

Your voice just gets transcribed into the box and submitted. Shorter batch is still batch, because the protocol is the part that did not advance. The protocol is the part we've had to learn. We just gave it a flattering name.

We call it prompt engineering and treat it like it's a power user skill: strip the label off, and it's a set of rules for packaging up good old batch. For example, tell it to think step by step. Give it examples.

Ask it to be an expert. Don't ask it to be an expert. Don't ask it that way. Paste more context. Paste less context. Only talk to it through markdown documents. We trade incantations. We've learned the magic words. That's the illusion.

It feels like mastery. But it's the same sort of mastery a punch card operator had: knowing exactly how to assemble the deck so the job wouldn't fail. Moreover, we've gotten good at prompting, or these black boxes, and that's the part that should bother us, not reassure us.

None of this means prompts are bad. Punch cards weren't bad. Command lines aren't bad. They're brilliant solutions for constraints at their time. But that's the whole question: is batch still theright protocol? Are we still prepackaging our intent for a machine that no longer needs us to?

Because it shouldn't need us anymore. It can ask a follow-up. It can clarify mid-thought. It can notice it's missing something and say so. It should be human conversational. Sherry Turkle of MIT puts it very well: conversation is the most human and humanizing thing we do.

It's one of humanity's superpowers. The capacity to engage and think isright there. And yet, we're still making people submit the deck and wait for the run. Even the punch card inherited its protocol, in fact. Batch came from the weaving loom.

You set the whole pattern in advance, then ran the cloth. The punch card got reused on computers by default. We're still standing at the same moment again. AI could finally meet us in the middle of a thought, got handed a protocol of a loom.

That's what I mean. The prompt is still a punch card, not because of how you encode it. The encoding is powerful and awesome. Because when the LLM is allowed to engage, only after you've packaged a complete turn and submitted it.

And this is where the mismatch bites. Model capacity is shooting straight up. Reasoning, speech, vision, memory, planning, all curving upwards. The interface protocol, flat. Still a box. Still a Submit button. Still the human doing all the work around the LLM.

The human still decides what context matters, still remembers what to ask, still chooses the timing, still notices the ambiguity, still repairs the output, still has to carefully engineer a prompt. But the intelligence feels magical. It's the interface that still feels like work.

Mismatch10:42

Ted Johnson10:42

And when it feels like work, when the output's wrong, when the magic words don't land, people blame themselves. They decide they're bad at this. They're not specific enough. They don't get AI. I want to say as clearly as I can: it is not our fault.

We are not bad at using AI. We are being asked to operate a brand-new kind of intelligence through a protocol of a punch card. The mismatch isn't the user. It's the interface. In the race to enable AI, we shortcut the interface.

Okay, let's make this concrete and familiar. A few weeks ago, my co-founder was using a frontier company's voice mode. These are known as speech-to-speech models. He asked it a normal question: "When is the next Timberwolves game?" Fine. It answered it quickly.

Then he pretended I showed up as if to speak to me and said, "Hey, Ted, come on in." He wasn't talking to the AI. But these models have no way to know that. So the AI did the only thing a prompt box can do: it took his speech as a turn and answered it.

"Sure, I'm here. What's on your mind?" That's not a good answer, but it's not a dumb model. It answered the first question perfectly. But it's a protocol with exactly one slot: your message, then its reply. It has no concept of who's speaking, whether the words were even meant for it.

Convergence12:06

Ted Johnson12:06

And the frontier companies want to make strides as well. OpenAI released GPT real-time too in late May and started trying it for their voice mode as well. It backchannels now. It goes, "Mm-hmm," and "Write." The little sounds we make to show we're listening actively.

We're seeing the field is converging on the same conclusion we built our company on: the interface has to stop being batch and start participating. Others are working on real-time conversation as well. This is NVIDIA's PersonaPlex, a research model, not ours.

Watch what happens when it gets interrupted.

Guest12:44

I've been thinking about starting a diet.

Ted Johnson12:46

Yeah, starting a diet can feel a bit daunting. But you could keep it simple. Focus on eating more veggies and fruits. Try to.

Guest12:54

Oh, before I forget, I signed up for a marathon.

Ted Johnson12:56

Allright, congrats on signing up for the marathon. That's a big challenge. You've got a lot of time. Focus on building a solid base with regular long runs. Stay hydrated. Make sure you fuelright before and after. And don't forget to stretch and take care of your feet.

PersonaPlex stops. It yields. It picks the thread back up. That's real turn-taking, listening and speaking at once, in real time.

Guest13:21

You need to come visit me. Because then we can go into the city.

Ted Johnson13:24

Oh, okay.

Guest13:25

Because that's the thing. Like, there's like the random spray paint, but then there's also like, I'm not sure. People must commission them. Like these massive mural spray paint pieces.

Ted Johnson13:37

Yeah, I think they do. And PersonaPlex's backchannels listens and lands where a person's would. Beyond conversational flow, there are lots of challenges, and it's a complex problem. Making listening noises is not really the same as knowing who's in the room.

These are not trained to tell that, "Hey, Ted" wasn't meant for it. But we are. We're working on improving the protocol to the models by giving it a better understanding of human and group conversation.

Guest 214:11

Good afternoon, everyone.

Group Conversation14:11

Guest 314:13

Good afternoon, Sam. Hi, Jordan.

Guest 414:17

Good afternoon. Good to see you both.

Guest 314:20

Quick one. Which requirement is this? Do we have an ID?

Guest 214:24

This is REQ-442, expense approvals.

Ted Johnson14:29

There. It answered a question. It only takes actions based on a utility-driven model. So it creates goals to fulfill, as it labels each of the participants' statements as a question, a proposal, an answer, and then only takes a turn when no one else is speaking or holding the floor.

Guest 314:48

Right. We need users to approve requests faster.

Guest 214:51

Yeah. The approval flow's too slow.

Guest 414:56

What kind of requests, though?

Guest 314:59

Expense approvals first. Access requests eventually.

Guest 215:10

AI, hold that. Actually, let's pause. Expense approvals or a general approval workflow?

Guest 315:19

Expense approvals. First release.

Guest 215:21

Access requests are future scope.

Guest 415:26

That changes the data model. Good to know.

Guest 215:31

AI, pull that up for everyone.

Ted Johnson15:33

Tracking determining who's the speaker's referring to is critical. In this case, it was easy with a direct reference to the AI. But it will happen again without a direct reference.

Guest 415:44

Oh, I'd forgotten that was a rule.

Guest 315:47

So over 5,000 needs a second approver?

Guest 215:50

Yep. Manager plus finance.

Guest 415:53

So a big one can't be a single tap.

Guest 215:57

Right. Over the limit, it routes to a second approver.

Guest 316:01

Agreed. Under 5, one tap's fine.

Guest 416:04

Works for me.

Guest 216:11

Okay. Agreed. Expense approvals, 5,000 threshold.

Ted Johnson16:17

Right there, the AI resolved the scope objective. No one wrote the prompt. No one packaged the turn and hit Submit. The system was in the conversation, following it, understanding, and choosing its moment.

Guest 216:31

AI, capture that for us. First release supports expense approvals only. Access requests are out of scope. Managers can approve or reject an expenseright from a notification. And per the finance controls policy, anything over $5,000 routes to a second approver.

Guest 316:59

Actually, make the threshold 10,000, not 5.

Guest 217:03

Want me to update the requirement to a 10,000 threshold? Yes.

Guest 417:10

AI, is this room free after the meeting?

Guest 217:13

Let me check. The room looks free after this.

Guest 417:15

Until 3?

Guest 217:15

But. Yes, it's yours until 3 o'clock.

Ted Johnson17:20

And that's the difference between a smart machine behind the same old prompt and an interface that finally participates. Here's the mindset shift I want to leave you with. AI is not just an intelligence technology. It's increasingly becoming an interface technology.

Mindset Shift17:36

Ted Johnson17:36

And if so, then book smart models alone are not enough. Stop picturing AI as a smarter machine hiding behind prompts, agents, loops, and all the old paradigms. We have to start seeing intelligence itself as a thing that can finally remove interface constraints and amplify human potential.

For 75 years, humans adapted to the machine. Its syntax, its forms, its timing, its batch. A system that can reason, listen, infer, adapt, should be able to meet us partway, if not all the way, instead. If AI is for users, then we should obsess about maximizing the interface.

So then the design question changes. It needs to become: what burden are we still putting on humans only because the machine used to be too limited to carry that burden itself? Ask that question, and the whole interface space opens up.

The answer isn't always chat. It isn't always voice, and not a wall of markdown. It's definitely not a decade-old set of digital constructs. Theright answer is the affordance humans already use with each other: communication, a question, a pause, a sketch, a checklist, a quiet aside, or saying nothing at all.

An interface where timing and modality aren't the human's job anymore, where choosing theright channel at theright moment is done by the AI. And as a usability person, this is the part that excites me the most. When you take that burden off people, the friction disappears and adoption follows.

Computing has mostly been about, to date, improving how humans encode their intent for machines. The punch card, type a command, click a menu, use your thumb on an iPhone, write a prompt. Every step was progress, and every step carried the old constraint forward into the next era.

A translation tax, a precision tax, context tax, repair tax. AI is our chance to put those down. Not by making everything magical. Not by making everything voice. Not by replacing human judgment. But by making computers, for once, more fluent with us.

Most talks and videos cover how to use or adopt AI. The deeper question is how AI intelligence changes the interface. Human conversation is the most human thing we do. Because if a machine can finally understand more of what we mean, then we can and should stop reshaping ourselves to be understood by it.

Thank you.