Should AI run on your own hardware? · Part 2 of 3 · 2 min read
Test local AI before buying the machine.
Design a small pilot with representative tasks, measurable quality, realistic concurrency, and explicit data boundaries.
Test the work you actually need done.
- Choose one task. Start with something bounded, such as locating the source for an internal policy answer or drafting a summary for review.
- Build a small test set. Include ten or more representative questions, difficult cases, and questions the documents cannot answer. Record the expected evidence.
- Use approved test data. Start with non-sensitive or appropriately prepared material. Get authorization before introducing restricted documents.
- Compare the full workflow. Record answer accuracy, source support, latency, and checking time. Compare with the current process and any acceptable hosted alternative.
- Test permissions. Give two test users different access. Confirm that answers and citations do not expose documents the requesting user cannot read.
Agree the pass criteria before the demonstration. A fluent answer is not enough. An unsupported answer should be marked as such, and a question outside the evidence should produce a clear limit rather than an invented response.
Three useful pilot questions.
These are possible starting points, not measured client outcomes.
- Operations: can staff find an approved procedure faster, with a link to the current version?
- Research: can a reviewer locate relevant passages across an authorized archive without losing attribution or context?
- Administration: can a team draft a first-pass summary of approved records while reducing the combined drafting and correction time?
Keep consequential decisions with a qualified person. Do not use a successful document-search demo as evidence that the same system is ready for legal advice, medical decisions, or autonomous payments.
For each pilot, name a user, an input, a useful output, and a measurable acceptance test. If the proposed benefit is only “we will have AI,” the use case needs more work.
Leave room to change the model.
Keep the application, permission rules, test set, and documents separate from the model where practical. That makes a later replacement easier to evaluate.
A newer model is not a guaranteed improvement for your workload. Before an upgrade, rerun the same tests, review its license, check resource requirements, and keep a rollback route. Measure Spanish-language performance directly if that is how your team works.
Document the result in one page: what passed, what failed, expected operating cost, and whether to proceed, revise, or stop. A pilot that rules out an unsuitable purchase has done useful work.
Edited September 29, 2026. Research and source dates are retained; this edit is not a fresh review of every statistic or legal development.
