Put AI to work responsibly · Part 3 of 4 · 3 min read
Measure AI results before you scale them.
Research and documented failures show what to test: accuracy, review time, permissions, and the cost of a wrong answer.
Research gives you test questions, not promised savings.
Studies measure particular tasks, people, tools, and periods. Use them to design a comparison in your own workflow, not to attach a universal productivity percentage to an AI purchase.
Customer support: a measured gain.
The 2025 published version of Generative AI at Work studied 5,172 customer-support agents and reported a 15% average increase in issues resolved per hour. Effects varied substantially across workers. That supports testing assistance in a defined support workflow; it does not establish the same gain for every role.
Experienced developers: a different result, and a later update.
METR's early-2025 randomized trial involved 16 experienced open-source developers and 246 tasks in familiar repositories. AI access increased completion time by 19% in that setting. The researchers explicitly warned against generalizing to all development work.
In a February 2026 update, METR reported that selection effects and measurement problems made its newer experiment an unreliable estimate of current productivity. It considered greater speedups likely, but could not confidently quantify them. The old slowdown figure should not be presented as today's universal result.
What this means for your pilot
Choose comparable tasks, define acceptable quality, and include checking and correction in the time measurement. Keep both successful and unsuccessful tasks in the record. Note the tool version and the people's experience so a later comparison means something.
Test both the answer and the access.
An assistant can fail by producing an unsupported answer or by exposing information through the systems connected to it. Those failures need different tests.
For answers, ask a reviewer to check the underlying source, not merely whether the output contains a citation. Include missing information, conflicting documents, and questions the system should decline. Record the consequences of a wrong answer and who can catch it before use.
For integrations, inventory tokens, permissions, and data recipients. Google Threat Intelligence's August 2025 Salesloft Drift investigation describes attackers using compromised OAuth tokens to access customer Salesforce data. The incident illustrates integration risk; it is not evidence that the underlying model invented an answer or that Salesforce itself was breached.
Test how access is granted, limited, monitored, and revoked. Start with read-only access where possible. Require appropriate authorization before the assistant can publish, send, delete, or make another consequential change.
Make the approved route practical.
A tool ban does not tell you whether people have stopped using the tool. Ask where AI already appears in the workflow and what problem people are trying to solve. Give them a way to request an approved alternative.
For the pilot, record four things: task quality, total completion time including review, errors with meaningful consequences, and user experience. Compare against the current process using similar work.
Scale when the results justify it. Keep the test set and repeat the comparison after a material model or workflow change. Research from another organization helps frame the experiment; it does not establish your expected savings.
Edited September 29, 2026. Research and source dates are retained; this edit is not a fresh review of every statistic or legal development.
