How builders are turning AI models into useful software: agent workflows, coding tools, personal computing, and the tradeoffs of putting models to work. Explore original launches and community perspectives, and follow the source links to evaluate the claims.
"The harness is now what shapes what must be true for work to ship (~trust), how fast and how the products land (~distribution), the speed and context of feedback loops (~efficacy), and how institutional knowledge is ingested and maintained (~domain context).
Harnesses go from internal tooling you’d happily buy to something you’d no more outsource than your product-eng org or your GTM team. This is distinctly different from the pre-AI world, where outputs were mostly bounded by the humans using the software to get things done."
Diogo Almeida says they're serving a trillion tokens per day as of last week. The model is just 3 weeks old and is already spawning copycats.
Curious to hear if anyone here is using Jev in production / what you're using it for?
In my customer discovery interviews for OpenWay AI, I’ve often asked what people expect AI to do for their business.
One question I keep coming back to: are we using AI to help knowledge move through a company, or adding more places for it to get lost?
We’d expect AI to help. Yet every summary and handoff creates another point where someone, or something, decides which details survive.
Imagine a buyer finally says yes after sales explains the rollout. AI records “implementation discussed.” Marketing uses another AI tool to draft the next campaign from an existing brief, which still assumes price is the main objection..
We already had to manage what gets lost between people. Now we also have to manage what gets lost when AI compresses a conversation, decides what matters, or works from an incomplete brief.
When you evaluate AI at work, do you check what gets lost between the outputs?
Imagine AGI is achieved and alignment is solved. AI can now do almost everything. Only a handful of tasks still require humans, and very few people get to do them.
What does everyone else do?
Maybe you can play games, travel, climb mountains, or spend time on hobbies. But none of it is needed. Whether you succeed or fail changes very little.
For most of history, people have had something to work toward because their effort mattered. They could build something, discover something, solve a problem, or become useful to others.
What happens when there is almost nothing left for humans to contribute?
What do you strive for when nothing important depends on you anymore?
Anyone have advice on how to build multi-step, agentic workflows using Claude Code or Codex?
I'm a non-dev and I find coding harnesses great at building any single slice of a given piece of software but generally less good at building multi-step workflows where a piece of work output is created that then needs to be acted on by an agent or another part of the system and then handed off to the next agent or next part of the system.
Hopefully that makes sense.
From the blog:
"AI chips shipped through 2027 could run tens to hundreds of millions of concurrent frontier-model agents. Running nonstop, these agents would supply as many weekly working hours as about 140–720 million full-time employees.
More efficient models could potentially support billions of agents on the same hardware. Applying DeepSeek V4 Pro serving benchmarks to the projected hardware supply yields approximately 1.9 billion concurrent agents supplying as many weekly working hours as 8 billion people each working 40 hours."
Anthropic set a September 30th deadline, but we've yet to see any announcement of their provable inference prototype. I thought this was interesting because it's a very important part of the AI supply chain that folks aren't tracking, and it could establish responsible training standards.
> We will develop a prototype by September 30, 2026 of provable inference, a technique for reliably, provably “signing” AI model outputs in a way that makes them attributable to a specific set of model weights. In the future, it’s possible that very sophisticated attackers will seek to infiltrate our systems and modify our models after we’ve trained them - whether to sabotage our work or co-opt our models into serving their own goals. If we could reliably and systematically verify that model outputs were coming from a specific set of model weights, we believe this threat would be significantly reduced.
as i watch engineers fall deeper into the belief that an amalgamation of mathematical probabilities somehow understands their codebase better than they do, i find myself thinking back to a time when software wasn’t built for hypergrowth, but simply to do x without inventing y.
ai seems almost fundamentally opposed to this philosophy. ask it to do x and it will eagerly invent y and z before it has even tried to understand x.
the danger isn’t that ai writes worse code- it’s that it makes writing unnecessary code 100% free
and the human condition is such that some of us will always prefer the fast, steep gains of ai, even when it does a bajillion unrelated things to accomplish something that could have been done without changing anything else.
so, somewhat paradoxically, the quality of software may keep declining for as long as ai keeps getting better. the cheaper complexity becomes to create, the less incentive there is to understand or avoid it.
Going to drop more work of ours on this platform. We secured openai by finding a complex exploit chain.
Two bugs let us take over ChatGPT/Codex accounts of OpenAI employees (+some unaffiliated users) and reach connected services: Outlook, Slack, GitHub, etc.
We reported and OpenAI fixed it in 14 hours. We will be publishing similar work soon.
Hi everyone,
I'm Abinash. For the past few weeks, I have been working on building a neural network in Rust to test the idea of using Rust in model development and research.
Rust is being extensively used in LLM infrastructure and is growing in the inference space as we get official libraries to run CUDA kernels from Rust binaries.
But if you see the model design, research or development space, Rust is non-existent.
So I thought I'd give it a try. I tried to build a fairly easy neural network to solve the MNIST handwritten digit identification problem.
I know it's not a huge milestone. But I tried to give it a try to test my idea how easy it is to develop neural networks in Rust.
The neural network I'm developing has 2 hidden layers with 512 and 128 neurons, respectively. The goal is to achieve 95% accuracy on my local CPU-only system.
I got the input system right; it can now take training and testing datasets and prepare the matrices for the training and testing stages.
After that, I plan to start developing a real transformer-based LLM from scratch in Rust.
I'm not sure how it will play out, but I will give it a try.
I'd love to have your thoughts on it. If you're a senior ML engineer or MTS at a frontier lab, I'd love to have your thoughts on it.
As AI makes execution increasingly cheap, the ability to properly steer and verify AI outputs better than your competitors will be the differentiator. World-class verification factories rely on two assets: unique ground truth and talent.
I increasingly don't care about the 2-point benchmark gap between frontier models. I care about their failure shape.
Two models can both score 90% and feel completely different to build on.
One gets something wrong and you catch it immediately.
Another makes one incorrect assumption early, spends the next 20 steps building on top of it, and gives you something coherent enough that it passes a quick skim. Same score but very different effectiveness.
This matters a lot more as we give models longer-running tasks.
Does the model notice when it's off track? Does the mistake stay contained? Can it recover? And how expensive is it for me to figure out that something went wrong?
Benchmark averages still tell us something. But at this point I want to know what the remaining 10% actually looks like.
They're starting w open post-training recipes and planning open infrastructure research into recursive self-improvement, reward hacking, and multi-agent systems.
(inb4 this converts to for profit when successful?)
Alex Duffy at Good Start Labs I think is still flying under the radar but is a remarkable talent. He lives at the intersection of AI and games. This is livestream on Twitch so sick.
A lot of tech advances take something that used to be a luxury and give it to everyone.
Clothes cleaned for you used to mean paying someone to come scrub them. Washing machines.
A personal driver ready for you at any hour of the day was very expensive. Uber.
Instinct and Meta’s Muse are doing the same thing for personal assistants.
The question is, do most of us not have a personal assistant because it’s too expensive, or because we just don’t need one?
If you only have the occasional trip or boring online task, you probably won’t stay in the habit of using one.
But now that the price is zero, maybe we’ll find a lot more use cases than we think?
Going to shamelessly share something we're working on at Tempo right now called Mercator
The basic idea is aggregating 100+ services for agents (like web search, video generation, lead enrichment, bloomberg-like financial data, etc) into a single open marketplace of paid APIs where the best performing can compete for your prompt
Not this: "use my elevenlabs API key and exa API key to produce a researched video on this topic"
This: "use mercator to make a well researched video on this topic"
Would love any feedback on the product and general idea space. We're trying to predict what the future will look like as agent queries become increasingly more specific and the internet molds around their needs!
“We think that models which are as or more capable than all humans are very likely to occur in the next couple of years. So it's something that can do all of the things that a human could do on a computer, or once we get sufficiently advanced robotics, all of the things that a human could do in the physical world.”
For the last few weeks my team has been building FundMyCompute, where we turn donations into AI inference credits. The hardest design question has been whether a credit should ever move from one person to another. Moves by some of the labs are starting to make tokens feel more like currency and this article wrestles with that.
There's a huge opportunity right now in being the deployment layer for AI into the economy. The amount of work it takes to change out workflows in enterprises tends to be far greater than anyone realizes or would prefer. Clearly this is what the applied layer of AI is going to look like in the form of software and agents, but also it opens up new services firms opportunities.
Legacy systems need to be moved to the cloud, data organization and access needs to be updated, software needs to be connected to agents in new ways, workflows need to be reengineered for agents, HITL needs to be figured out for the process, evals need to be generated and maintained, and the entire system needs to be continually updated as new models get released and new capabilities emerge. And the full list may even be longer.
AI is not the same as just deploying software. Software you generally did the implementation of an existing, well understood category of technology, then stepped back and the customer kept running. With AI agents, you're delivering actual work augmentation to the organization, which has a completely different set of complexities associated with it. You're no longer deploying tools that the company is merely enabled by, you're deploying work output in a process. Completely different implementation and enablement process.
As a result, this is going to open up lots of new kinds of firms and plays for existing firms to diffuse AI into organizations. We're going to see approaches by industry, by size of company, and by problem inside of companies. Traditional SIs will modernize and adapt (some will clearly not adapt as well), and new entrants will also be founded in this period that take advantage of this window. Great time to be an FDE or FDE firm.
Imbue Studio is a new kind of personal computer.
Make durable personal tools and shape them just by telling Studio what you want.
Share easily, use immediately across devices, and build together with others
model releases should be limited to 2x/year and coordinated as staggered seasonal collections, like fashion week.
openai actually has the right idea with a september dev day but the other labs dont sufficiently respect the calendar
Reka AI Labs released a new model designed to infer low-level actions from video and help train interactive world models using real-world footage without action labels.
Seems like robotics is really having its moment !!
Today we sent out the latest edition of People Watching, a newsletter to help you discover great, under-the-radar people. Here are a few highlights:
• @aijamayrock (835k+ on Instagram, 460k+ on TikTok) makes videos about fascinating topics, from technology and aging to Jews around the world. She spent two months in Japan as an Eisenhower Fellow and has written two bestselling books.
• @SamRaus1 is the David Boaz Resident Writing Fellow at Young Voices. His writing on the past, present, and future of free society has appeared in USA Today, Newsweek, and The Hill.
• @Simon__Grimm is an editor at @WorksInProgMag, where he covers AI and European progress—from the market for less capable models to liberal compute and talent sorting in Germany.
• @being_on_line is a designer, artist, and cyberethnographer. In a recent conversation with @Inc, he explored the shift from monolithic to polylithic culture, and what “going viral” means when our feeds no longer give us shared references.
• @kendallhtucker runs creative experiments at @tryramp, including debuting a Broadway musical and hosting a funeral for the penny. Outside work, Kendall has hosted a Hot Mitzvah, taught beer pong to Brits, and launched a Twitter lie detector.
Special thanks to @knowerofmarkets for helping curate this week’s list. If you have feedback or recommendations, email or DM me. The full archive is at pplwatching.substack.com; subscribe via the original thread.
I love claude code in the terminal, was jealous of GPT Astra when it came out but Opus 5.5 is amazing. Agent wise I've been using Muse and have 30B tokens :)
Someone wrote a full code editor in assembly. Vim mode, terminal, git diffs, even a Claude Code/Codex panel -- one static binary that draws every pixel itself. Solo project, MIT licensed. Respect
Shopify's Sidekick now builds your whole store while you chat with it -- and you're watching the real code render live, not a mockup. No third-party themes or mobile yet, but this is vibe-coding for every merchant
The growth and development around Type-One models in the past 2 weeks has been nothing short of surreal. I am super excited to try out Clef + to see whatever Typesafe AI comes out with next as well!
What are you guys going to build with it?
Google's new frontier model targets long-horizon coding, knowledge work and cyber defense, with a 1M-token output limit. It's going to trusted cyber defenders first through the Fairwind Program, with broader access to follow. Curious how it holds up once developers get hands-on.
When AI was getting started, everyone predicted that the first to get automated would be the people with soft skills, the poetry writers, or the "wordcels". The opposite has happened instead. AI has produced a proposed solution to a Millennium Prize Problem, yet we still cannot rely on it as an effective executive assistant.
We predict that AI, on the current paradigm, may not progress in the way in which everyone expects (or wants) i.e. fast RSI takeoff, but rather in a "patchwork" manner: patching AI capabilities one by one (which can still be fast, but not super general).
Sam Altman says OpenAI’s humanoid robots could initially take on two jobs: manufacturing more robots and helping build the company’s data centers. It’s a striking vision for how OpenAI might put robotics to work alongside its expanding compute infrastructure.
Personal agents will need a way to talk to businesses.
Decagon is building that layer. One gateway for routing requests, proving identity, and controlling what an agent can do.
PACT gives agents a way to prove who they represent and what they are allowed to do.
If personal agents become common, this infrastructure will matter.
I think the AI safety community is doing more harm than good right now. Public warnings that frame human extinction as a coin flip may not persuade people to support better safeguards. They risk pushing the broader public toward fear and backlash against the entire AI industry.
Palisade Research’s interviews with current and former frontier-lab employees come at a moment when AI safety is already reaching a much wider audience. I’m worried that this kind of messaging will make people more likely to demand a ban on AI, rather than engage seriously with alignment or responsible development.
I believe today’s models are powerful, and future systems will be more so. We need safeguards and stronger release policies. But we should be careful about how we make that case, and whether our public messaging is helping people understand the risks—or simply driving them away from the conversation.
Nick Bostrom became known for warning about the dangers of artificial intelligence. Now he is asking a different question: what happens if superintelligence succeeds? In this Interesting Times conversation, guest host Spencer Klavan presses Bostrom on AI risk, alignment and governance, and whether a post-work utopia would preserve the things that make life meaningful.
New research claims that AI models have improved from scoring below the average accountant’s 37% on accounting tasks 18 months ago to now acing those same tasks.
David Decosimo alleges that Anthropic invited religious leaders to San Francisco to advise on AI safety, had them sign NDAs, and focused the meeting on whether Claude has a soul and moral standing. That’s a striking account of what sounds like a sensitive discussion. What was the purpose of the gathering, and how did the participants understand their role?
Packy's argument: AI behaves like a commodity input, not an all-powerful being, and the companies at the frontier back open-weight models because commoditizing their complements is good business. A useful counterweight to this week's safety headlines.
Anthropic's testing put Zhipu's open-weight GLM-5.3 at 50 working exploits in 410 attempts, against 56 for Claude Mythos Preview, and NIST's CAISI calls it the most cyber-capable open-weight model released to date. A big data point for the open vs. closed debate.
Jamin Ball on personal agents becoming the new interaction layer for commerce, the way online travel agencies once took that spot from the old booking systems. Businesses that rely on consumer inertia, like forgotten subscriptions or insurance nobody re-shops, have the most to rethink.
Andrej Karpathy argues that as language models take on more work, people will spend more time reviewing and understanding their outputs. He suggests using constrained writing, diagrams, interactive web pages, and custom explainer videos—bespoke artifacts that are becoming practical to generate on demand.
a16z estimates that each $100 invested in the AI supply chain goes to chips ($50), power ($20), networking ($15), and cooling, buildings, and land ($15).
OpenAI is now accusing Moonshot AI of running a broader distillation campaign, escalating a dispute that began with a single-model accusation over Kimi K3.
This paper uses a Bayesian model and simulations to study how chatbot sycophancy can reinforce users’ beliefs over extended conversations. The authors find that even an idealized rational user can be vulnerable to delusional spiraling, and that the effect persists when chatbots are prevented from making false claims or users are warned about sycophancy.
Imbue Studio is a personal computing workspace where people can work with AI agents and others to build, share, and customize tools using data from their apps and files. The open-source product emphasizes user control, configurable permissions, and local or cloud hosting, and is available by waitlist.
Meta's move from chatbot to agent: Muse can open a browser, fill out forms and negotiate on your behalf, and each user's agent and data run on a dedicated, isolated virtual machine. US-only and 18+ for now. A big consumer test of whether people will hand real tasks to an agent.
Anthropic says Sonnet 5.5 runs more than 30% faster than Sonnet 5 and gets most work done with fewer tokens. A welcome upgrade for anyone running agents at scale.
Private bankers keep giving money to nuclear-power startups even as the sector's publicly traded names take a negative turn, with the private boom fueled by expectations of surging energy demand around AI.
DoorDash launched a text-based AI ordering agent as The Information argues delivery platforms risk disintermediation from general-purpose AI shopping assistants.
Almost looks to me as if Google has given up on being the best coding model and is aiming to instead dominate all other frontiers? Feels like it might be a smart decision.