Harness Engineering as UX Design
I'm working on an article on everything we learned to make cfo.ai's Ari agent as capable as we can. What would you be most curious about?
Read original source ↗I'm working on an article on everything we learned to make cfo.ai's Ari agent as capable as we can. What would you be most curious about?
Read original source ↗I have noticed that current agents struggle a lot with spreadsheets, why? It's somewhat surprising that Anthropic or OpenAI can't make this very good easily.
HOOOO boy. That is the most interesting thing we've learned while building Ari and is the single thing that made Ari possible. In short - the models aren't the problem. The problem is in the format of a spreadsheet and context window limits. It is a distinctly different shape than code or writing.
Seperate but related; back when I was at Morgan Stanley they rolled out copilot for all employees and the older members of my team did not want to use it at all (I think an even split of them not believing it was capable, them being scared of it becoming capable, and not wanting to fit it into existing workflows). As you think about making your agent more capable, how are you also thinking about making this something that CFOs/finance teams trained in legacy still want to use?
How do you make sure that the harness is optimized for prompt caching and how do you monitor it in product + monitor agents performance in the wild.
Have you experimented with giving the model both representations? Like, extract the spreadsheet structurally, but also render the relevant sheet/range as an image and let the model use the screenshot to infer the visual hierarchy and produce a better semantic representation of the cells? Seems like a sweet spot would be structured cell data + a visual rendering, then have the model reconcile the two Maybe even tile/crop aggressively the screenshot of the spreadsheet. Obviously macros/VBA are a separate problem entirely
Also a great question. I think we've seen a similar story already in coding agents. Pre December 2025, the quality of the coding agents weren't very good, and a lot of engineers avoided using them. I think that's basically completely changed in the last 9 months. So my view is that when it becomes capable, people will start using it. And if it's not capable, people rationally avoid using it.
Will definitely write extensively about this. But tl;dr the harness is the easy part, the biggest lever is in tool design.
Yes, definitely - it's a good theory. In practice we didn't find visuals to improve outcomes significantly. The problem is deeper than not being able to see visually. Will be writing about what we learned in depth.
Curious to know whether you had challenges with the AI enforcing responses formatted in a particular way and how you worked through those. I gave up on getting LLMs to comply with a particular output format years ago, so I'm curious what SOTAs are capable of in practice.
Love the framing. i'd be most curious about how you distinguish changes that make Ari genuinely more capable from those that merely make it feel more capable. With the underlying model held constant, which changes to context, tools, memory, verification, and failure recovery produced the biggest measurable lift? Also curious how you designed the two interfaces at once: Ari’s interface to the world, and the human’s interface to understand, correct, and trust Ari. i am building a harness for early stage founders and my biggest dilemma is that founders love a warm, encouraging copilot (engagement goes through the roof) but their outcomes deteriorate. Whereas if the harness pushes them to do hard things, the user's outcomes improve but the overall engagement goes down.
Supportive of this take, many decisions during harness development are absolutely in the domain of UX design
What's the 10x or 100x better experience than excel, and what's the required ttv to convince someone in finance the agent is that 10x or 100x experience?
what’s your approach to helping Ari self organize?
Sign in to join the discussion.