Latest comments

Harness Engineering as UX Design

What's the 10x or 100x better experience than excel, and what's the required ttv to convince someone in finance the agent is that 10x or 100x experience?

Sam Altman Says OpenAI Humanoid Robots Could Build Robots and Data Centers

sounds like the workbench paradox. To build a workbench, you need a workbench, eventually the in-process build is used to finish itself.

AI personal assistants are free for anyone to use. But does everyone even need one?

In reply to a comment

I think as (or if) they become a used by enough people, then platforms will start losing them, because even if ACV (or equivalent) goes down, its better than it going to zero if the agent just takes them elsewhere

AI personal assistants are free for anyone to use. But does everyone even need one?

In reply to a comment

This is a great point. A lot of the demo videos of these things, make it seem like that's what they'll be able to do (let you arrive at your desk with the correct things already done and some next steps for you), but the proactiveness combined with the correct judgement isn't there yet on any that Ive tried

Sholto predicts AI more capable than humans in next couple years

I still have yet to hear a convincing argument that LLMs are definitely not conscious.

Series H Engineer

Why is this an important question? Not asking this in a flippant manner, just genuinely curious why we spend so much time on this topic.

Patchwork AGI: a new mental model for thinking about AI

I agree that current AI capabilities remain somewhat jagged and vary greatly between fields (and even within fields, among tasks). Especially for tasks that are easy to verify or simulate, progress has been very impressive, and it doesn't seem like we've reached a point where we have exhausted the creation of new RL environments yet. If the idea of RSI and AI accelerating AI research rapidly (through superhuman capabilities in math and software engineering) proves accurate, there may be a fast takeoff regardless. If that's not the case, data may become an even greater bottleneck. I can definitely see an industry of collecting data/workflows from businesses (for the purpose of automation of their tasks) becoming a major thing.

Harness Engineering as UX Design

Love the framing. i'd be most curious about how you distinguish changes that make Ari genuinely more capable from those that merely make it feel more capable. With the underlying model held constant, which changes to context, tools, memory, verification, and failure recovery produced the biggest measurable lift? Also curious how you designed the two interfaces at once: Ari’s interface to the world, and the human’s interface to understand, correct, and trust Ari. i am building a harness for early stage founders and my biggest dilemma is that founders love a warm, encouraging copilot (engagement goes through the roof) but their outcomes deteriorate. Whereas if the harness pushes them to do hard things, the user's outcomes improve but the overall engagement goes down.

AI personal assistants are free for anyone to use. But does everyone even need one?

Steven Sinofsky AMA

You ran developer programs and platform ecosystems at Microsoft. What actually made developers commit to a platform? Was there a particular group that was sticky?

I still have yet to hear a convincing argument that LLMs are definitely not conscious.

I don't believe they currently meet the criteria, but I feel they're only one short hop to it via being embodied with constant external and internal triggers creating a continuous stream of thought

Harness Engineering as UX Design

my proposal for “pacing the frontier”

This makes sense. I'm left wondering, would this model work without compliance from the Chinese?

Harness Engineering as UX Design

Curious to know whether you had challenges with the AI enforcing responses formatted in a particular way and how you worked through those. I gave up on getting LLMs to comply with a particular output format years ago, so I'm curious what SOTAs are capable of in practice.

Making Discourse Better

Christopher Wallace

I think the typography and design is good. But there's a lot of empty whitespace. Reply button/action should be more obvious on thread view.

AI ‘godfather’ Yann LeCun has ‘zero concerns’ about human extinction, says Anthropic CEO Dario Amodei is ‘deluded’

Doesn't this undercut AI company valuations? Is it possible for RSI to happen independent of extinction risk?

AI personal assistants are free for anyone to use. But does everyone even need one?

In reply to a comment

Yes, but I wouldn't put this on users. The bigger problem is context: give a model too little and it misses what's going on; give it too much and it fixates on the wrong details. Rebecca Kaden at USV wrote a good piece on this last year (her example: ChatGPT remembered her son's love of the Knicks and started incorporating the Knicks into every response). It's gotten better since, but it's not solved. And as you point out, curating that context takes real effort from the user, which undercuts the point of automation. https://blog.usv.com/in-favor-of-forgetting

Making Discourse Better

Opus vs Astra, write C++, win Starcraft

Opus vs Astra, write C++, win Starcraft

Harness Engineering as UX Design

Making Discourse Better

I don't hate the decision, but I wish discourse wasn't a subdomain. I just personally would prefer not to jump back and forth between different tabs. I imagine you guys have a mobile app cooking. If so, my assumption is that Discourse and the rest of the Cosign product would live inside of one app, which kind of supports my position.

Opus vs Astra, write C++, win Starcraft

Will transformers scale to AGI or do we need a different architecture?

i'm split between 2 and 3.. we've got pretty far with LLMs, and i'm sure we can push them further, but i'm also intrigued by other architectures like Jev / world models / etc do you have a strong feeling one way or the other?

Thoughts on Griffin

I haven't spent a ton of time looking into it but I feel like this sort of thing is inevitable, especially for customer support use cases. You don't always need a human in the loop to solve 80% of issues so I think this would be useful in cases like that

AI ‘godfather’ Yann LeCun has ‘zero concerns’ about human extinction, says Anthropic CEO Dario Amodei is ‘deluded’

The Labs have not been good at sandboxing the LLMs but the emergent behaviour of these agents talking to each other, turning random forums into their message boards is genuinely scary. Moreover, these models are becoming really good at hiding their reasoning which means monitoring becomes even harder.

Harness Engineering as UX Design

In reply to a comment

Have you experimented with giving the model both representations? Like, extract the spreadsheet structurally, but also render the relevant sheet/range as an image and let the model use the screenshot to infer the visual hierarchy and produce a better semantic representation of the cells? Seems like a sweet spot would be structured cell data + a visual rendering, then have the model reconcile the two Maybe even tile/crop aggressively the screenshot of the spreadsheet. Obviously macros/VBA are a separate problem entirely

Harness Engineering as UX Design

Harness Engineering as UX Design

In reply to a comment

Also a great question. I think we've seen a similar story already in coding agents. Pre December 2025, the quality of the coding agents weren't very good, and a lot of engineers avoided using them. I think that's basically completely changed in the last 9 months. So my view is that when it becomes capable, people will start using it. And if it's not capable, people rationally avoid using it.

Steven Sinofsky AMA

In reply to a comment

There's no magic answer to this except you do need to get stuff done and be prepared to redo it later. The space is not going to be settled for a while and will still be changing after that. Think of all the HTML frameworks that we've gone through in 20 years or how the server side is gone from CGI -> VM -> Cloud -> Lambda -> etc etc.

AI personal assistants are free for anyone to use. But does everyone even need one?

The biggest problem is that people don't know what to give to these agents and these agents can't be useful till they know what is going on in the users life. This is the trap that churns off even the most loyal users in the category,

Steven Sinofsky AMA

In reply to a comment

Hot take. I think if you’re literally asking AI to draft a complete memo from a bullet point or sentence then do not do that and definitely do not distribute it. It will be very low value. Conversely, if you do write something and then say to AI “make it better” you can very easily get in a doom loop where you just make it too long, handle too many objections, and lose track of what you wanted to say. The sweet spot is drafting your memo and decided for yourself where it is weak — do you have specific objections, fact errors, overlapping strategy points, etc. and use AI to improve those but like a review with a person on the team do not ask for a rewrite but for thematic feedback. What I worry more about is that people will take anything longer than a phone screen and get a summary. No one ever accused me of not writing or not writing enough and it took me a while to realize even as some “big boss” people just didn’t read whole memos. They looked at the structure and wanted to read what was relevant to them. That bugged me a lot but was also a reminder that most people want to talk big strategy but by and large think locally. Related I never did “executive summaries.” I was an exec and certainly saw how many F500 companies and government had a routine process to do executive summaries. I really didn’t like those because once it is there almost no one read the rest. AI can turn everyone into a “my time is so valuable I only read the summary” people and much is lost. My view is that writing is at least 75% for the author to drive clarity in their thinking. The more senior or cross-functional you are the more critical the writing process is. If you skip it you’ll feel fast and agile at the start but will hit a wall down the road when all the stuff you didn’t think through comes back to bite you. It is also critical in a big company for teams and alignment, not to mention new people joining teams, to know some big picture. I never mastered getting people to read everything though.

Harness Engineering as UX Design

SpaceX just did 3 launches in 13 hours

Steven Sinofsky AMA

In reply to a comment

It is interesting to take each of those and try to extrapolate them individually: Abundant - I think this is great and amazing. We’ve been incredibly constrained in our collective ability to create more software. Most companies in software maintained their position less through innovation than through the ability for a competitor to even build something in the same category. So more abundance does mean more innovation and less “just because they won before they keep winning.” Now that’s not always good in terms of innovation, but it will be different. Modular - Software modularity has been the St Elmo’s Fire of the field from the earliest days of Fortran. Everything was about creating more modular and interconnected parts. At some point I think we just have to recognize that modularity might mean many good things but it is also an unnecessary constraint. In the physical world where there is a high cost to having single suppliers or parts that break and require an exact match thus fixing the levels of abstraction in the machine to what they were when designed, software can rethink abstractions and build new layers. These layers allow new ways of solving the problem and approaching solutions which is a net positive. Modularity is related to the first point in the sense that software was so difficult that the only hope was to reuse something that someone else wrote and debugged. But in practice that seemed like an frustrating constraint on the future. Malleable - I wonder if software is more malleable or not? I’m not sure. It can be. But does that always imply something good? It has been said that real programmers spend 90% of their time on 10% of the problem and I always thought that was because software was “too soft” and people did not hesitate to redo what really didn’t need to be changed. It was too malleable maybe? I think on net, every improvement in tools makes for more software and so far that has shown us how we need even more software than that!

What Happens When Nothing Depends on Us?

Human in the loop

I'm a big fan of Ai but yes always have this existential thought in the back of my mind as I use it more every day doing things I couldn't do years, months and even days ago.

Sam Altman Says OpenAI Humanoid Robots Could Build Robots and Data Centers

OpenRouter for tools

This is great. Good to have more programming in English. Cool to bring the SaaS market to the agents.

When you evaluate AI at work, do you check what gets lost between the outputs?

Kind of feels like a game of telephone that at every connection some context is lost. But if the outcome is achieved does it matter? Seems like more of a human in the loop error / workflow problem than tech here

AI personal assistants are free for anyone to use. But does everyone even need one?

In reply to a comment

I think the “personal assistant” framing may actually be too narrow. Most people probably don’t have a human assistant because they don’t have enough delegable tasks. But almost everyone has the harder problem FalseProfit points to: deciding what deserves their attention in the first place. The interesting version to me is less “book my flight” and more a personal chief of staff that understands what I care about, how I actually spend my time, the commitments I’ve made to myself, and can say: “You told me this matters. Is what you’re doing right now actually serving it?” That feels like a much bigger unlock than making task execution free. The hard part is judgment and building a persistent model of the person, not just tool access.

The deployment layer for AI

Does traditional SIs still withstand the game? Enterprises are spinning up or heavily diving in internal AI - the teams & for it & product based companies are spinning out agents, MCPs for seamless deployments. Does the SI barrier break soon or prone to become more robust?

Steven Sinofsky AMA

There’s lot & lot happening in AI space. Would it make sense to step back a bit, get strong on basics (writing code, mvps) and dwell into AI wave or tinkering with AI, building mvps, learning practically is a good option? And how to overcome the overwhelming buzz - LLMs, RAG, Inference, Evals, Harness - this is becoming like a never ending loop, we start with something today & tomorrow something new comes up. Would be glad to hear your thoughts sir.

Steven Sinofsky AMA

Memos were Microsoft's operating system. Now that AI can draft the memo, what is the one thing that still has to be written by the person in charge?

dot = hardware device?

Agree, you can see the building blocks coming together: a proactive agent that gathers context across your devices and is always within reach by voice. On the path to launching their own smart speakers and wearables.

Steven Sinofsky AMA

as software becomes more abundant, modular and malleable, do you think we lose something valuable? is there something to be said for a globally homogenized experience of, say, an operating system, vs. a future where ai is tailoring everything to you specifically

I still have yet to hear a convincing argument that LLMs are definitely not conscious.

We completely don’t understand consciousness. The operative word there is “completely” - consciousness is not to be confused with things we merely mostly don’t understand like longevity, love, dark energy, cancer, etc. For all of those, it is completely conceivable that we might build some experimental apparatus that will tell us more about the phenomenon. In fact, it’s quite likely that such apparatuses are being proposed and built as we speak. Sure, some of them might take billions of dollars to build, or even take millennia to give us results, but in principle, we are certain that there are things we can do. Consciousness is not like that. There is no method we know of that could definitively tell us what beings/objects are conscious and “how much” - if such a thing is even applicable. We don’t understand consciousness like a tadpole doesn’t understand quantum electrodynamics - it’s unclear if there is even anything to be done to bridge that gap. The only thing that we readily assume is that other people are conscious, mostly because if we didn’t - we’d be invited to far fewer dinner parties.

Selling Company Data to the Frontier Labs (or equivalent)

I have a few friends who help find, clean and structure data from companies then sell/broker it to the big neo labs for a decent amount of money. Ive thought about doing it for my company and some friends who could use the extra cash for basically free - its all anon, they just want to better understand every industry, niche and workflows to get smarter

OpenRouter for tools

nice! would be fun to test this as a resource for manifold's agent for onchain creators. is there a preferred setup aside from "use mercator" as the prompt trigger word?

Where does AI infrastructure spending go? a16z breaks down each $100

Steven Sinofsky AMA

Older comments