AIAI EngineerNov 2, 2023· 18:56

The Intelligent Interface: Sam Whitmore & Jason Yuan of New Computer

Sam Whitmore and Jason Yuan of New Computer advocate for AI interfaces that adapt to human context by using implicit signals like proximity, gesture, and tone, rather than forcing users to adapt to rigid input modes. They demonstrate a pose-detection system that switches between keyboard and voice input based on user distance from the screen. They explore social gestures for rejecting calls and propose new physics metaphors like blending images by shaking an iPad. They argue that probabilistic materials like generative AI require familiar metaphors from nature and human behavior, and that mixing input and output modalities in real time yields more natural experiences. The demos, all speculative fictions, point toward future hardware where context-aware reasoning becomes the norm.

Transcript

Intro0:00

Sam Whitmore0:15

Hi everybody, thanks for having us here today. Um, we're super excited to be here. I'm Sam, and I'm one of the co-founders of New Computer.

Jason Yuan0:23

And I'm Jason, the other co-founder. And we're really excited that we are starting today by letting you all see our paws up close, which is amazing. Um, so, you know, when Sam and I started New Computer, we— we did so because we believed that for so long we've taken certain metaphors and abstractions and tools for granted.

And for the first time in what feels like 40 years, we can finally change all of that and we can start thinking from first principles. What our relationship, not only with computing but with intelligence, period, should look like in the future.

So what do we mean by intelligence? Because, uh, you know, sometimes I'm on the internet and I wonder if it even exists. Um, well, one way to think about intelligence is, uh, the ability to sort of take in lots of information, different types, different volumes from different sources, um, visualized as dots here, and sort of find ways to make sense of it all, find ways to reason, find ways to find meaning.

Intelligence Defined1:00

Jason Yuan1:27

Um, and as human beings, as carbon-based life forms, we do this through, uh, a process where at first we use our senses to sort of perceive the world around us. Um, then we, you know, process that information in our heads and then, given what we think, we then choose a reaction.

Um, so if we're lucky, we are blessed with at least five senses. Six, when I've had four margaritas. Um, but as humans, we sort of are in- inherently capable of just processing all of this at the same time, and then that actually is how our short-term memory gets to work.

Um, and taking in all of this context and information, we then get to form what's called a theory of mind. Um, what is going on? What is, you know, how is the world relating to meright now? What should I be doing about it?

So we sense, we think, and then we react. Um, and how do we react? Well, um, there's a lot of thingsright now. Uh, but if we take it all the way back to the Stone Age and we think real simple, um, a lot of how people used to react and communicate is just unintelligible grunts.

Um, and then one day we— that sort of evolved into, uh, language as we know it. Um, and to this day, that's still something that we rely on to communicate and react to the world around us. Um, and that's also how a lot of us think.

So we have language. Um, but the language of communication is so much broader than just language. We're standing here on stageright now. I'm making eye contact with some of you. Nice shirt. Um, and I'm making gestures. I'm wearing these ridiculous gloves.

I'm looking at Sam. I'm looking at things. I'm pointing at things. Um, and I can hear sort of laughter or I can hear people, you know, thinking. I'm taking in lots of information at once, andright now I'm sensing, thinking, and reacting.

Sam Whitmore3:27

So this year, um, well, last year technically, we saw a really amazing thing happen, um, kind of with the advent of ChatGPT, I would say, where we saw the beginnings of a computer start to approximate that same loop where input was coming in in the form of language.

Language Loop3:27

Sam Whitmore3:44

There was some reasoning process, um, however that, uh, actually works. Um, and then the output felt also like language coming back to us. And this was very inspiring to me and Jason, and we've been spending a lot of time this past year thinking about what's next and how this gets to feel even more natural, um, for people to interact with computers specifically.

And so today we wanted to take you on a tour of a few demos. Um, one, um, which you can do with a computerright now, um, and then a few which are kind of with futuristic or, uh, next-generation hardware which may be available soon.

And knowing that you're all engineers, we know that this will kind of get the sparks flowing, um, the ideas flowing for seeing how, like, you might use, um, some of these things that are coming out soon or things that exist today to build things that feel more natural.

Adaptive Input4:38

Sam Whitmore4:38

So I'll start by getting to a demo, and I will say, um, this is a live audio-visual demo, so I am foolish enough to make that choice. So we will see how it goes.

Jason Yuan4:53

Um, before we show any demos, it's prudent to point out that none of these represent the product we are building. They are simply.

Sam Whitmore5:02

Yes.

Jason Yuan5:03

Pieces, stories, inspiration.

Sam Whitmore5:06

So the point of this first demo is to imagine we have a lot of things where we're saying like, "Okay, is text theright input? Is audio theright input?" And we've been thinking about

it's not if those are theright things, but when. So in this case, you'll see some measurements happening on the left here. What's actually happening is that this has access to my camera, and it's taking, uh, real-time pose measurements of where I am with relative to the screen.

So I just— it knows I'm at the keyboard basically because it's making that assessment, and you can see the reasoning in the side here where it's saying, "User is close to screen, we'll use keyboard input. User is facing screen, we'll use text output."

And so this— we're using an LLM to actually make that choice as it— as it goes to the response. So let's try something else. And again, demo gods be nice 'cause this may not work at all. Um, but if I now walk away and it doesn't detect me anymore, it should now actually start listening to me.

Hello. Can you hear me? Are you gonna respond?

Jason Yuan6:09

I think that's a no.

Sam Whitmore6:11

It might not respond. But basically what we are attempting to build here is, like, if I want to actually talk to the computer in a really natural way, um, if I'm there next to the keyboard, I should not— it should not be paying attention to my, uh, voice or any sounds, ambient sounds.

And if I walk away from the keyboard, I might wanna have a conversation with it, like, walk around the room. It is listening. It seems to not— to decided not to actually talk back, but, oh, it's talking.

Is there something you need help on sounds like an interesting project, Samantha? How is your talk going so far?

Sam Whitmore6:54

Yay.

Yes. You can see it paid attention and it decided to ignore me for a while.

But anyway, this is— this is just like a toy demo. You can see here we have, um, this is how it's working kind of behind the scenes. It's, like, trying to decide if I'm close to the keyboard, facing the screen, not facing the screen, and use that all as inputs to decide whether it should talk to me or, um, just display the text as on the interface.

Um, cool. So.

Jason Yuan7:33

The reason why we— we think this is interesting is because we think, you know, people are naturally sensitive to other people. And, um, we— we think computers, instead of asking people to adapt to computers to be, like, come up to me and type and whatever, should find ways to try to adapt to circumstances and contexts of people.

Sam Whitmore7:56

Exactly. So, um, again here, it's like in this case, it's adapting to where I am by using the pose detection, whether or not I'm actually in the process of talking to it to decide to update its own world state, use an LLM to actually do that, and then use the LLM to respond using the knowledge of that world state.

And so this is a really simple and, as you can see, kind of hacky demo that is what something you could build today. In theory, you could imagine how this could be, like, a really cool native way to, uh, interact with an LLM on your computer where you don't have to worry about the input finality at all.

Um, so again, takeaways are consider, like, the explicit inputs, what I'm typing, what I'm saying, along with implicit, where I am. Um, there's other things you could do with that, like tone and emotion detection. Um, you could plug in a whole bunch of different signals that you wanna extract from that.

Jason Yuan8:46

And you can even imagine if I'm in the frame with Sam and the agent knows Sam and she had recently been complaining about me, she'd probably not bring that up until I leave the frame. Um.

Sam Whitmore8:57

Yep. And as we mentioned that, um, using it as a reasoning engine and then next one. Cool. And yeah, and then we're adapting. So we wanna get to the futuristic stuff. Um, Jason has been spending a lot of time imagining this, so he's gonna walk you through a few things that might exist shortly in the near future when new hardware comes out.

Jason Yuan9:17

So, um, when we think future, we still think the sensing, thinking, react loop will, will take place. To preface all of this, these are my personal speculative fictions, not representative of anything that I think might actually happen. Um, and this is a very conservative view of the next 1 to 12 months maybe.

Social Interface9:17

Jason Yuan9:38

So it's not a true future, future AGI god worshiping type situation. Um, so let's start with, uh, what I call, like, a social interface. Um, we're all really excited about, you know, certain headsets being released at certain points.

Um, and one thing that I think is interesting about some headsets is they have sensors and they have hand tracking and eye tracking. Um, and just like how I'm being expressiveright now, maybe there comes a day where I can be such with a computer that sort of lives with me.

So here I'm— here I am in my apartment minding my own business. Um, and my ex decides to, uh, FaceTime me. Um,

and now I've declined the call. You know, with historically deterministic interfaces, um, I would've had to, like, find the hang up button or go, like, "Hey, Alexa, decline call." Like, thinking commands, thinking computers speak. But, like, as a person, I can be like, "Fuck off."

You know, I can be like, "I'm busy." I can be like, "I'm sick." You know, like, all of this stuff, the computer should be able to interpret for me and, you know, send, send, uh, what's his name again?

Tox-toxictrashee-ist, whatever, on his merry way. Um, so explicit social gestures can be a great way to determine user user intent, like the way I just showed now. Um, but we should also consider interpreting implicit gestures. If I give a really fast gesture with a slow gesture, my mood, my tone, how far away I am.

Um, but we should also be conscious of social cultural norms. Different gestures mean different things in different societies, and it might mean, you know, as you scale your application or hardware to different locales, this is something that you should pay attention to.

Generative Physics11:19

Jason Yuan11:19

Now, I wanna move on to talk about what I call new physics, and this part is super fun. Um, this demo is based on, um, a little, uh, I think on an iPad, which, you know, has over five daily active users in the world.

It's very popular. Um, and here I'm imagining, like, okay, Midjourney. If I was the founder of Midjourney, I would be putting all my resources and making some sort of, uh, Midjourney Canvas app for iPad. So in this one, I've asked Midjourney to create, uh, Balenciaga Naruto, which now I'm realizing kind of looks like me.

Um, so let's think about the iPad. It's like this big slab that you can, like, touch and fiddle with,right? So what do I wanna do? Okay, I wanna, like, edit this photo. Um, but first I need to make space.

How do I do that? Well, very easy. You just, you know, um, you can just zoom out, and now you have extra space. Very obvious. We do this all the time. Um, I kind of think my cat would look really good in that outfit.

So I kind of wanna find a way to do that here. Let me just ask AI real quick. Um, hey, random AI, send me pictures of my cat. And, you know, the AI knows me and has context and gives me pictures of my cat.

And then what do I do here? Well, why can't we just

take one of the photos and sort of just blend them with the other? Um, and the metaphor you're seeing here as you sort of work with these photos, they start glowing when you pick them up. And what does light how you guys know the Pink Floyd, uh, Dark Side of the Moon album cover?

Like, we're really familiar with the idea that light can sort of, uh, provide different colors and, and sort of concentrate back into one form. And we're leaning into that metaphor here implicitly. Um, and so it's now created something that looks 50% human, 50% cat, 100% cringe.

I don't really like this. How do we remix this? What is a gesture? What is the thing we do in real life that's remixing? Um, for me, it's a margarita, and for Sam, it's her morning fuel. We shake a blender bottle.

So why, why can't we work with intelligent materials the same way that we work with real materials and just blend it up? This is totally doableright now. David, why aren't you building this? If you don't build this, I'm gonna build this.

It's fine. Um, so, you know, here the metaphor is like, what we're trying to say is, you know, think about familiar universal metaphors like physics, like light, like metabols, like squishy, like fog, whatever. Because, you know, if you're designing an iPhone, you have to be very cognizant of the qualities of alu-aluminium and titanium to make an iPhone.

But generative intelligence is a probabilistic material that's sort of more fluid. Maybe it's fog, maybe it's mercury. Um, and, you know, for this reason, maybe metaphors that are really rigid, like wood or paper or metal, aren't theright metaphors to use for some of these experiences.

Um, so finally, we wanna walk you through an experience that's inherently mixed modal, um, slash mixed reality. Um, let's imagine for a second there's a piece of hardware coming out that's a wearable that has a camera on it and has a microphone and can maybe project things.

Mixed Reality14:15

Jason Yuan14:31

I don't know if such a thing will ever exist, but let's imagine for a second it does. Um, I'm sort of browsing this book, this Beyoncé tour book, and I see these images that I find really inspiring. Um, what I'm trying to do here is what if I could just point at something on my desk and say, like, "This is cool," and have the sort of device, uh, pick up on that and, and, and indicate that it's heard me and it's gonna do something by sort of projection mapping the sort of feedback.

Um, this is, you know, this demo doesn't really have sound, but the way this would work is ideally a combination of voice and gesture at the same time. Um, and obviously this gesture is really easy to make mistakes with.

So anytime you work with probabilistic materials, you wanna provide a graceful way out. So in this case, I've accidentally tapped this photo. Why can't I just flick it away like dust and be like, "That's wrong." I don't wanna press an undo button.

I don't wanna press Command Z. I just wanna flick it away. Um, really leaning into the physics of it. Um, so now that I found two pieces, I'm kind of like, okay, I wanna send this to two of my friends who, hmm, there was a friend who I said I would do Halloween with, but I can't remember their name.

Um, what do I do here? I should ask AI. I should be like, "Who is that friend I said I'd spend Halloween with?" And you'll notice here that, like, we're imagining sort of projection mapped UI pieces that can work with the context of the world you're inright now, such that you don't have to go fish out a phone or use cumbersome voice commands.

Um, it just all sorts of naturally melds in with the world. Um, and, you know, crucially, I think one point we want to make is voice in doesn't need to mean voice out. Gesture in doesn't need to mean gesture out.

And visual UI in does not need to mean visual UI out. We can mix these modalities in real time for whatever makes sense in whatever context you're in. So given that interactions that require multiple simultaneous inputs are now possible, um, it's our job as designers and developers to sort of think on behalf of the user and think what's the appropriate output given the current context and be smart about it.

Um, yeah.

Sam Whitmore16:44

Yeah. So again, the takeaways, as we mentioned, it's this idea of we have a lot of sensors and, and contextual modalities available to us as ingredients, even today. There'll be more tomorrow, as you kind of saw with these upcoming, uh, potential hardware releases.

Takeaways16:44

Sam Whitmore16:58

Um, but even now with a laptop, with things like typing speed, with things like, uh, the tone of voice, there's a lot of ways that you could gather context and extract signals from it. You could choose to process it in a variety of different ways.

And so all of that can now be passed to an LLM and used in a reasoning layer, which decides how, um, both to respond in words and also how to present that information. Um, and so basically everything can now be an input, and your output could be everywhere and have every format.

Um, at the same time, one might say everything everywhere all at once.

Jason Yuan17:37

Well, you wanna be intentional with it. You know, you if someone wants to generate a photo on their Apple Watch, you're like, "Why, why?" Like, no, use your freaking phone. Jesus. Um, anyway, and the last thing we'll say is, um, probabilistic interfaces are hard because they have lots of different outputs.

So a really great way to sort of ground these interfaces is to lean into familiar metaphors, whether they are from nature, from physics, or even from human-made tools and materials, like buttons for now. Um, and, you know, social norms is also a material that we work with,right?

So your banking AI agent probably shouldn't be able to have a deep philosophical chat with you. That just socially doesn't make sense.

Sam Whitmore18:16

That would feel weird.

Jason Yuan18:17

Exactly. Um, but on the same note, we've, we've, we've related all of these interfaces to what humans perceive and experience now. But what might a truly intelligent interface look like in the future where if we think we are where we areright now is pseudomorphism, what is the abstraction layer above that?

And that's kind of for us to figure out. Um, so with that, um, yeah, I think that's all.

Sam Whitmore18:45

Thank you.