What Products Are Worth When AI Agents Become Buyers
How AI agents could change software pricing, product comparisons, and the way we define value
A company is renewing its customer-support software. Today, choosing a replacement could take weeks: reading reviews, sitting through demonstrations, comparing plans, negotiating a contract, and running a pilot. Even then, the buyer is comparing claims made under different conditions.
Now imagine giving an authorized AI agent a representative set of support requests, a budget, security requirements, and permission to try several products. It could measure which answers were correct, how often a person had to step in, how long the work took, and what the final bill would have been. The recommendation would come from the company’s own workload, not only from a vendor’s presentation.
At first, this looks like faster procurement. The deeper change is that someone must first decide what a good result means. Accuracy might matter more than speed. A clean handoff to a person might be worth more than an automatic answer. A lower price might be unacceptable if it creates greater risk. An agent can run the test, but it cannot avoid the question behind it: what does this buyer value?
When buyers can test before they trust
People use brands, reviews, analyst reports, and sales conversations because testing every product is expensive. Reputation lets a buyer borrow other people’s experience before investing time and money.
Agents could change the order of that process when products offer clear digital interfaces. G2 is already experimenting with this idea through an AI Agent Evaluations section that publishes independent performance benchmarks run by Swept AI. G2 describes the methodology as experimental and currently limited to customer-service agents. This is not evidence that agents have replaced software buyers. It is evidence that a major software marketplace is beginning to supplement reviews with direct tests of performance.
Reviews would still matter. A short test can measure accuracy, speed, and cost, but it cannot reveal how a vendor behaves during a difficult implementation, whether employees will trust the product after six months, or what happens during an outage. The likely shift is not from reviews to measurement. It is toward a market in which each covers the other’s blind spots.
For product companies, that changes what it means to be discoverable. Reputation may earn a place in the test. The product then has to make its value visible through actual use.
A price contains a theory of value
Business software is often sold by the seat because a buyer can count how many employees need access. That model becomes less natural when an agent moves through several services to finish one job. Paying for a full month of every service can cost more than the work itself.
Usage pricing moves the bill closer to the work. Stripe’s usage-based billing can meter requests, processing time, storage, or tokens. Stripe provides the measuring equipment, but the seller still chooses the unit. A request, a minute, and a token describe activity. None is automatically the same thing as value.
Outcome pricing tries to close that gap by charging for a result. Intercom publishes outcome-based prices for its Fin AI agent, including resolutions and completed handoffs. The attraction is clear: if the system does not produce a useful result, the customer should not pay as though it did.
The difficult part is deciding what happened. Intercom may assume a problem was resolved when a customer does not ask for more help. The answer may have worked, or the customer may have given up. Seats value access. Usage values activity. Outcomes value a defined change in the world. Each pricing unit is also a claim about what the product is for.
Agents may make these claims easier to compare by estimating the total cost of completing a job. That creates pressure for clearer prices, but not necessarily lower ones. Vendors can keep the savings, marketplaces can collect fees, and sellers can hide complexity inside minimums or bundles. Competition lowers prices most reliably when the buyer can actually choose a substitute.
Agents may weaken the old forms of lock-in
Leaving a product can require moving years of data, rebuilding integrations, retraining employees, passing another security review, and risking disruption. Agents do not erase those costs, but they can reduce some of them by translating data, recreating routine connections, testing replacements, and coordinating a migration in smaller steps.
They may also let a company choose one product for each part of a workflow instead of accepting an entire ecosystem simply because its data already lives there. The UK Competition and Markets Authority connects data portability and easy switching with meaningful choice for this reason. When data is portable and interfaces are standard, ownership of the historical record becomes a weaker moat. Products at the center of a company will remain harder to replace, but capturing a market through data lock-in alone may become increasingly difficult.
This could also change what a product is. An agent might use one service for research, another for verification, and a third for a specialized action. Products would compete as capabilities that can be combined, not only as complete ecosystems that must be adopted whole.
The scorecard becomes a shared skill
An agent can ask a narrower question than a software ranking: which product produced the best result for this workflow, under our rules, at this total cost? That sounds objective, but it does not remove judgment. It relocates judgment into the scorecard.
A test can reward whatever is easiest to count. If it measures only speed, vendors may sacrifice qualities that slow the work down. If it rewards questions answered without a person, healthy escalation begins to look like failure. If it measures only immediate cost, it can miss employee confidence, customer trust, or a problem that appears months later.
The metrics themselves will change. Seat count, response time, tickets closed, and cost per transaction are useful partly because they are easy to count. They are proxies. As agents observe more of a process and its downstream effects, evaluations can become richer and closer to the outcomes people actually value. That may change both the apparent ideal and the products optimized to reach it.
The important change is not that an AI agent replaces a procurement department. Procurement itself can become a reusable capability: a method for discovering, testing, comparing, and governing purchases that improves as organizations contribute experience. If that work can be expressed through language, policies, examples, and evaluations, its accumulated lessons can travel farther than any one department.
The same may become true of customer support, compliance, property management, and other forms of operational judgment. Systems can compare how the same function is performed across organizations and carry forward what works. Some differences that once looked like strategic choices may turn out to be duplicated inefficiency.
Convergence does not mean every company makes the same decision. Goals, risks, relationships, and constraints remain local. Someone still has to decide what should be optimized, which exceptions matter, and when a metric has stopped representing the outcome it was meant to approximate.
A transaction still needs human authority
For an agent to buy responsibly, several questions have to be answered. Who is it acting for? What is it allowed to do? Did the product satisfy the buyer’s test? How much should be paid, and what record should remain?
Identity establishes authority. Evaluation supplies evidence. Payment moves money under agreed rules. The record makes the decision reviewable later. The previous essays in this series examined machine-readable payments and identity for AI agents separately. Auth.md, MCP, MPP, and x402 provide pieces of this larger system, but none decides whether a product is good, whether its price is fair, or whether the agent should be allowed to buy it.
Once agents direct meaningful spending, whoever designs a popular evaluation gains market power. A vendor can tune for the benchmark. A marketplace can decide which products are visible and which qualities count. Automated markets can become more efficient while also creating new gatekeepers.
A healthier market would keep those choices contestable. Buyers need portable data, clear bills, spending limits, and evaluations they can inspect. Sellers need a fair opportunity to be tested. People need the ability to revise what “best” means when the scorecard fails them.
The future value of a product may therefore depend on more than what it claims to offer. Products will have to show what happened, what it cost, and why the result should count. That could make value more concrete and prices more honest. It also makes the definition of value too important to leave unquestioned.