Skip to content
HN On Hacker News ↗

What Would A Serious AI Product Look Like?

▲ 176 points • 84 comments • by lumpa • 2w ago • HN discussion ↗

Pangram verdict · v3.3

We believe that this entire text is human-written.

0 %

AI likelihood · overall

Human
100% human-written 0% AI-generated
SEGMENTS · HUMAN 1 of 1
SEGMENTS · AI 0 of 1
WORD COUNT 1,712
PEAK AI % 0% · §1
Analyzed
Sep 28
backend: pangram/v3.3
Segments scanned
1 windows
avg 1712 words each
Distribution
100 / 0%
human / AI fraction
Verdict
Human
Pangram v3.3

Article text · 1,712 words · 1 segments analyzed

Human AI-generated
§1 Human · 0%

One of the issues that I have with the current generation of “AI” products is that they do not appear to take their own premises seriously. I look at a plethora of obsequious chatbots claiming to be serious tools for problem solving, and I think, this is not what a problem-solving tool would look like. Even before we get to the tremendous ethical problems with the frontier labs, it is this impression of their composition as a product that makes me feel, constantly, whenever I am interacting with them, that they are less a software product than that they are a grift, a scam designed to make me feel like I am interacting with a product that has capabilities that it simply does not, to try to lull me into a false sense of security that I can trust it. The frontier labs are of course the worst offenders, but every criticism here applies just as much to Ollama, which (if anything, due to the obviously poorer quality of the available models themselves) needs these features even more than the frontier labs do. Here, I will set down a few features that might convince me that an LLM-based product, particularly one focused on research or software development, was actually serious about helping me do useful things with it. Make “Checking For Mistakes” A First-Class Feature This is the biggest issue, and the major reason that I was inspired to write this post. It is a truth universally acknowledged, that AIs cannot reliably provide information. I could cite a ton of news articles and studies about this fact, but there is no need. Every single chatbot admits this, up front, in a fine-print disclaimer as a core part of their user interface. Gemini says “AI can make mistakes, so double-check responses”, Claude says “Claude is AI and can make mistakes. Please double-check responses.1” ChatGPT says “ChatGPT can make mistakes. Check important info.”. Every time I see that last one, I wonder how I’m supposed to know what “info” is supposed to be “important”. All of these warnings are all small, gray text, painfully obviously included as legalese to push responsibility back onto the user rather than to help with anything. This is a core limitation of all these products. Checking their output is a part of the workflow for using them that: you absolutely cannot skip or skimp on without creating risks to yourself and whoever you are conveying its output to, and, it is very easy to skip or skimp on and you are encouraged at every turn to do so, because “just trust the output” is one of the quickest ways to save time. A chatbot product that took this weakness seriously, as an actual consideration for using it, would put a checkbox next to every claim in its output. It would be a 2-column worksheet, where you’ve got the LLM output in the first column, and next to it, human notes in the second column, explaining what work went into checking this claim, and a big checkbox that you would only check off after you believe you’d checked its claims thoroughly enough. Coding assistants would need to have some version of this as well. Right now, this is pushed off into code review, which means it is a dark pattern which subtly encourages the “author”2 to offload this work to their code reviewer without ever looking. Once again, “it’s probably fine, I don’t need to check” is the quickest way to save time and churn out those PRs faster. It might even be useful for coding harnesses to have some affordance for checking code before it even runs tests. As the vendors themselves have admitted, it’s not just expensive to burn tokens on your “AI”, you also end up burning far more compute on the AI. Being able to check your diffs before sending them over to uselessly exhaust your testing compute cluster would be useful. If your product tells me that it makes mistakes and I must be the one to check for the mistakes, but then gives me zero tools to check for mistakes, I cannot take it seriously. More Citations to Check, And More Details Most chatbots prefer to give an answer, rather than a citation. In my own personal use, I find that when asked to provide a list of citations with clearly marked sources for each one, they will appear to “get bored” halfway through the list and simply stop including citations at some point. When the bots include citations at all, present them as inline annotations that say nothing but the domain name of the search result, in a font so small that it’s barely legible, and an equally indecipherable icon that is fewer than 16 pixels on a side. This is backwards. Now, I am aware that these citations do come from somewhere, and in an attempt to reduce hallucinations, all of the major providers support some form of “grounding”3, and that those little barely-readable citation links are referencing actual structures in the RAG pipeline and not just potentially-hallucinated tokens, but I’m not talking about the underlying machinery in the model, I’m talking about the presentation to the user. Plus, regardless of whether a snippet of text came from a RAG query, we know that LLMs can never provide an authoritative result; it’s a fundamental limitation of the technology. They can still garble the results of RAG as much as they can misrepresent any other training data. This means that it must never present its results as authoritative. If you ask an AI to do research queries, every result should be presented as a list of citations. Moreover, the presentation should display each citation as a large object of in its own right, with clearly identified metadata, including not just the site where it was found but its publication date and, if possible, the name of the author. The literal, unmodified quotation (not from RAG, not a summary: a quotation extracted with a regular program and not an LLM) should be front-and-center, larger than any AI-generated text. If the AI product wants to editorialize or summarize (which should not always be necessary!), the AI-generated text should be presented as small text underneath the citation that has been found, de-emphasized as much as the disclaimer is right now, at the very least until the user has verified that the summary is accurate. Perhaps, for a research project, a “did you read the citation” checkbox might even be helpful. If your product openly tells me that it will scramble, misrepresent, or omit its citations in its summaries, and I must read the original human-authored citations to be sure, but then gives me no tools to track my reading of those citations or even any way to find them, I cannot take it seriously. No First-Person Output, No Apologies There is no reason for a software development or research tool to use first-person language to describe itself. They should not do so. In fact they should not be allowed to do so. There is also no reason that they should ever apologize. It is a waste of everyone’s time; it’s a waste for the chatbot to generate the apology, it’s a waste for the user to read the apology, and it’s a waste for the user to respond to the apology. Yet they unfailingly do this upon every correction. The vendors of these tools know that they are routinely causing mental-health crises. In response, they have added non-functional “guard rails” that can still, in 2026, easily be bypassed.4 A product seriously interested in helping with productivity would correct this glaringly obvious flaw, focus on the task at hand, and stop emitting useless verbiage. In the previous two sections, I tried to focus on ways in which the harness would be constructed differently even if the LLM technology is fundamentally impossible to improve; in this case, I have to assume that the labs have some control over the model itself. But unless they are truly incapable of influencing their output (and all their “benchmarks” and “capabilities” seem to indicate that they can control it very tightly) they ought to be building models that are much less verbose. More Non-Natural-Language User Interfaces Although natural language could hypothetically be a powerful interface for interacting with a computer system, the practical upshot of LLM natural language interfaces is that these interfaces are imprecise and repetitive, full of superstitions masquerading as “best practices”. The inputs are a mess and the resulting outputs are a mess. The general way of addressing this unstructured mess is to allow the chatbot to directly take action in response to the user’s input; in other words to supply it with “tools” via an MCP server. But again, this is backwards. If we cannot even express our intent clearly in the first place, why are we trusting this system to take potentially destructive and harmful actions on our behalf? Instead, I would expect a product that was seriously invested in helping me accomplish specific tasks, to have user interfaces specific to those tasks. Is it supposed to be able to be a security scanner that can discover OWASP top 10 bugs in a codebase? Have a button for that. Build that functionality into your harness, train it directly into the model, use smaller models that can satisfy that functionality more effectively than throwing it at the planet-sized brain of Fable or whatever. I’m aware that there are small software startups that do something like this, but they are bolted on to the side of the main model providers’ APIs, not integrated into the core of the product and not using their own models and AI systems to achieve consistent and repeatable results. Strong Data Provenance Indicators Chatbots produce data tables pulled from websites, from APIs, from MCP tools or from summarizing and scrambling the user’s input. In order to provide the illusion of a seamless interface, this data is presented in-line regardless of where it comes from. But some of these outputs are produced mechanically via regular old API calls, for example, from the result of calling a tool or querying a website, but presented uniformly. But there is a huge difference between an authoritative data source