Skip to main content

Market Daily

Can You Trust What AI Says Your Customers Think?

Can You Trust What AI Says Your Customers Think?
Photo Courtesy: Unsplash.com

A growing class of software promises something companies have wanted for decades. It reads everything customers say, the reviews and support tickets and app-store comments and social posts, and tells teams in plain language what it all adds up to. The timing helps the pitch. Survey response rates have slid for years, so the old metrics now rest on a thinner and thinner slice of customers, while large language models have made it practical to process hundreds of thousands of comments at once. InMoment and other established experience-management platforms are adding these capabilities to their suites, and a younger group of AI-native companies, including Unwrap, Chattermill, SentiSum, and unitQ, is pursuing the same goal. The category even has a name now: customer intelligence, and a fast-growing list of buyers.

The question the marketing tends to skip is whether the machine is reading the feedback correctly.

Where the Models Fall Short

Start with sentiment analysis, the layer that sits underneath most of these tools. It has well-documented blind spots. The ones researchers return to again and again are sarcasm and irony, negation, ambiguous words, and comments that carry more than one feeling at once. “Great, another delay” sounds positive on the surface but means the opposite. “I guess it works, if you have an hour to spare” is a complaint dressed as praise. A review that admires a product’s design and savages its price does not collapse cleanly into a single positive or negative score. Models have improved at all of this, but they still lean heavily on context. Accuracy measured on tidy test data tends to fall as it aligns with how people actually write, including typos, slang, half-sentences, and industry shorthand.

How far it falls is itself disputed. Independent benchmarking of automated sentiment tools shows accuracy varies widely by method and by dataset, and can drop sharply on hard, nuanced text, which is exactly the kind of text customer feedback is full of. A useful reality check is buried in how these systems are built. Even on the simple task of labeling a comment as positive or negative, studies have found that trained human reviewers agree only around 80 percent of the time, and agreement falls further on finer-grained judgments. If people cannot fully agree on what a comment means, a model trained on their labels inherits that ambiguity. That is an argument for judging these tools by how well they fit a specific use case, on a company’s own messy data, rather than by a vendor’s headline accuracy figure measured under ideal conditions.

The Newer Problem: Hallucination

Generative AI adds a new problem on top of the old one. The summaries and written narratives that make these products feel impressive are produced by language models that sometimes hallucinate, generating fluent, confident text that the underlying data does not actually support. Picture a weekly digest that announces customers love the new checkout flow, when the comments behind it were a mix of sarcasm and a vocal minority, or a summary that quietly promotes a handful of complaints into a top trend because they happened to be phrased in strong language. In a business setting, those errors are dangerous precisely because they are well written. A clean, confident paragraph is exactly the kind of output a busy executive is least likely to stop and question, and the wrong takeaway can move a roadmap. Across the AI field, reliability has moved from a technical footnote to something boards now ask about directly.

To be fair, the alternative is not perfect either. Surveys have their own reliability problems: leading questions, a narrow, self-selected group that actually responds, and scores that blend a customer’s feelings about price, product, and service into one number nobody can act on cleanly. So the comparison that matters is between two imperfect methods, each with its own failure modes.

How the Vendors Answer

The companies building these tools know all of this, and how they handle it is becoming a point of competition. Unwrap is one example. Founded in 2022 by Ryan Millner and Ashwin Singhania, two former Amazon Alexa product leaders, it spun out of the Allen Institute for AI’s startup incubator and now applies natural-language processing to unstructured feedback for product and customer-experience teams. It lists customers including Microsoft, DoorDash, Lyft and Perplexity. The company says its models tag feedback with a high degree of precision and that every insight can be traced back to the original customer comment. Buyers should still press on how any accuracy figure is measured, since precision and recall can be traded off and a single percentage rarely tells the whole story. But the design principle behind the claim is the right one. A tool earns trust to the degree that a person can click from a stated trend down to the raw comments underneath it and check the work, rather than taking a generated summary on faith.

That principle points to how careful teams actually use this software. They treat it as a first pass that turns an unreadable pile of text into something a human can review and judge. In practice, that means a few concrete habits. They keep insights grounded in and linked to the source comments, so any claim can be audited. They spot-check the model’s output against a sample they have read themselves, and re-check it when they move into a new product area or a new language where the model may be weaker. They keep a person in the loop on decisions that carry real cost. And they read AI-scored sentiment as a direction rather than a precise measure, as a smart manager would treat any single metric. Those habits are what separate a useful deployment from a credulous one.

The Real Trade

None of this is an argument against the tools. The alternative, leaving the vast majority of customer feedback unread or extrapolating from a survey that a handful of people bothered to answer, has obvious problems of its own. AI analysis genuinely surfaces issues that would otherwise remain buried until they surface in churn. It is better understood as a trade. These systems sacrifice some precision for enormous coverage. A survey hands you a relatively clean number from a narrow group. AI analysis hands you a noisier read across almost everything customers are saying, and it does it fast enough to catch a problem while it is still small.

Whether that trade pays off depends less on a vendor’s accuracy claims than on the buyer’s discipline. The companies getting real value tend to use these tools to decide which questions are worth a human’s attention, then send a human to answer them. They build a habit of verification and treat a confident summary as a hypothesis to test before acting on it. The technology keeps improving, and grounding insights in source text, exposing confidence levels, and keeping humans in the loop will all help. But none of it removes the basic safeguard: someone reading the actual comments before acting on what the software says they mean.

Market Daily

This article features branded content from a third party. Opinions in this article do not reflect the opinions and beliefs of Market Daily.