Support & Education For whoever signs the data agreement An idea we are exploring Field notes · August 2026

Your own model,
on your own desk.
Nothing leaves the room.

Most of what an organization does with AI starts with a paste. A case note, an intake form, a grant report, an interview transcript, a spreadsheet of families: somebody copies it into a window that belongs to a company somewhere else. The tool is excellent and the answer comes back in seconds, and the thing you pasted is now on a server you do not control, under terms nobody in the office has read. For a marketing team that is a policy question, and it can wait. For a foundation holding the case files of four hundred families, or the recorded voices of people who trusted an archive with their memories, it has to be settled before anyone starts.

There is a version of this where the model is a file, the file lives on one machine in your office, and nothing you type into it goes anywhere. No vendor receives it, and there is no redaction step somebody has to get right every time. The models are open weights anyone can download, and the machine is a desktop computer. These are notes on what that adds for the organizations we work with, what it can do today, why it gets better on its own, and the parts of it that are harder than the pitch.

Skip to what we would build with it ↓
Read this as a proposal.
The capabilities are what we can switch on now. The trajectory is our read of the last three years. When we have run this for a real team, this page gets their numbers.

What “private”
actually means here.

↳ the word gets used for three different things. this is the strict one.

Every AI vendor will tell you your data is private. What they mean varies. The consumer tier means it trains on what you paste unless you find the setting. The business tier means it does not train, and the retention is whatever the contract says, and the data still crosses the internet and sits on their machines while it is processed. Both of those are real forms of privacy and we have a whole page on choosing between them. This page is about the third thing, which works differently from both.

A model is, physically, a very large file of numbers. Running it means loading that file into memory and doing arithmetic on it. If the file is on a machine in your office and the arithmetic happens there, then the sentence you typed and the document you attached were never transmitted to anyone. There is no vendor involved, so there is no retention policy to read and no terms-of-service update to watch for. What makes it private is where the cable goes.

Where the data goesNowhere

The question and the document are read from your disk into your memory, and the answer is written back to your screen. Pull the network cable out of the machine and the assistant keeps working. You can run that test yourself, and no cloud tool passes it.

What the vendor seesThere is no vendor

The people who trained the model published the weights and walked away. They do not know you downloaded it, they do not know what you ask it, and they have no way to find out. That is what open weights means. There is no vendor contract to negotiate, so the licence is the document to read.

What it costs to askElectricity

There is no per-question charge and no per-seat plan. The cost of a question is the power the box draws while it thinks. Teams behave differently when that is true. Nobody rations questions, and nobody quietly falls back on the free consumer tool because the paid one hit its limit.

↳ the workaround this replaces, and why it was never good enough:

You stop scrubbing the data before you use it.

The standard advice for using a cloud model on sensitive material is to strip the sensitive part first. Replace names with placeholders, pseudonymise the identifiers, tokenise the account numbers, and send the model the hollowed-out version. An entire product category exists to do this, and organizations pay for it. It has two problems that never go away.

The first is that it has to be right every single time, and it is done by a rule or a small model that does not know that Doña Carmen from the second floor is a name. One miss and the record you were protecting is on the server anyway, and you do not know which one. The second problem is harder to see: the context you scrubbed out was the reason the question was worth asking. A navigator wants to know which of her families is at risk of a benefit cliff next month. Anonymise the families and there is no question left.

With the model in the room, you do not scrub. You ask the real question about the real record, in the words you would use to a colleague, and the record stays where it was.

The data does not go to the model. The model comes to the data.

What it can do
for you, today.

↳ what current open models do on one desk, with the caveats attached.

The open models crossed a line in the last year. The weights the laboratories publish for free are, for the everyday work of an organization, as good as the paid tools were eighteen months ago, and they run on a desktop computer that plugs into a wall. The hardware side is a purchase and not much more: one machine with as much memory as Apple will sell, in the closet where the router lives. The list below is what a team can do with it, with a note on how well each one works.

1

Read your documents and answer from them

Point it at a folder, a database or an archive and ask questions in plain Spanish or English. It answers from your material and cites where it found it, so you can check the answer against the page it came from.

works well. the citation is the part to insist on
2

Summarise, draft and translate, in both languages

A case note becomes a follow-up letter, a quarter of figures becomes a funder narrative, and a policy memo becomes the Spanish a family will read. Register and tone are set once, by you, and held from then on.

works well. review before it leaves the building
3

Pull fields out of forms, scans and photos

The current models can read pictures. An intake form photographed on a phone, a scanned letter from 1962, a family’s photograph with a note on the back: names, dates, places and amounts come out as structured fields for a person to confirm.

good, not perfect. a person confirms
4

Transcribe and tag recordings

Hours of interviews, meetings or voice notes become searchable text with speakers and timestamps, tagged by topic, without the audio going to a transcription service. Puerto Rican Spanish included.

works well. dialect and noise need a check
5

Sort, triage and flag

Applications by readiness, tickets by urgency, the participant who has gone quiet, the record that contradicts last quarter. It reads all of it so a person only has to read the ones that need a decision.

works well. the rule is yours, not the model’s
6

Do small pieces of work end to end

Given a connector and a rule, it can draft the letter, file it in the right folder and put the reminder on the calendar, then stop for approval. Longer chains of steps are where it still needs a person alongside.

early. short chains yes, long ones not yet
Real, and good enough

For the work above

The open models running on one desk today do all six of those at a level that would have been the best available in the world eighteen months ago. For a case file, a grant report or an archive, the gap to the most expensive cloud model rarely matters.

Real, unfixed

For the hardest reasoning

The largest closed models are still ahead on the hardest problems, and the very largest open ones do not fit on a desk. If a task needs that tier, the question goes out and the record stays home. We describe how that works below.

What we would build
with it, by partner.

↳ things we would build. none of these organizations has signed anything.

The organizations we build for have a lot in common: small teams, a great deal of trust placed in them by the people they serve, and files those people would not want in a stranger’s data center. Every one of them has a use for a model, and almost none of them can responsibly paste their real work into one. What follows is what changes when the model is in the building, written against work we already know.

1

Family Navigator: the caseload that can finally be asked a question

The navigators at Vimenti, PSI’s integrated family services center, carry families through benefits, housing, schooling and work, and the record of each family is the most sensitive thing in the office. Today a navigator who wants to know which of my families has a recertification due before October, and which of those also lost hours last month has to know it already, or build a spreadsheet. With the model reading the case files on a machine in Vimenti’s own office, she can type that in Spanish on a Tuesday and read the list. The follow-up letter, in the register the family expects, is the next thing she asks for. The files never leave the office, and the navigator never opens a browser tab to a vendor.

case files stay in the building
2

The benefit-cliff calculator, explained to the family it applies to

The Calculadora de Beneficios, built with the Instituto del Desarrollo de la Juventud and PSI, computes what a raise costs a family in lost assistance. The arithmetic is the easy part. The hard part is explaining, to a specific mother with a specific job offer, what the number means and what her options are, in plain Puerto Rican Spanish, without a caseworker in the room. A local model that has read the rules and her own figures can do that explanation, and can draft the testimony a policy analyst needs for a hearing next week from a hundred real cases, none of which had to be anonymised first.

real cases, not de-identified ones
3

Fundación Rimas: reading applications from minors without a vendor in the loop

La Academia’s participation platform collects applications, mentor notes and progress records about young people, and we have written before about what it means to have minors in a dataset. A model that helps staff triage applications, summarise a mentor’s notes before a check-in, or spot a participant who has gone quiet is useful. Sending a fifteen-year-old’s essay to a cloud vendor to get that help is a decision a board should have to make out loud. On a machine in the foundation’s office nobody has to make it.

nobody’s child in anyone’s training set
4

La Voz del Centro: asking a question across a thousand hours of voices

The archive holds decades of interviews with people who agreed to be recorded for a purpose. Transcribing them, tagging them and making them searchable is already work we do, and the processing stack is described on that page by what it does rather than who sells it. A model on the archive’s own machine takes the next step: find every time anyone describes the 1985 landslide, in their own words, or what did the interviewees born before 1940 say about the sugar mills, answered from the transcripts with citations back to the minute. The voices never go to a transcription vendor, which is what the interviewees were promised.

the voices stay with the archive
5

Archivo Negro: cataloguing family memory without exporting it

Families hand El Archivo Vivo photographs, letters and stories, and the curators contextualise and safeguard them. The bottleneck is description: every item needs a caption, a date estimate, names, places, a link to the collection it belongs in. A vision-capable local model can draft that description from the scan and the note the family sent, for a curator to correct, at the pace the submissions arrive. The family’s photograph of their grandmother is never uploaded to a company that will keep a copy, and the archive can put that in writing.

the grandmother’s photo stays in Puerto Rico
6

PSI’s front door: the grant report that writes its first draft from the data

A platform for social impact runs on reporting: to funders, to boards, to the public. The numbers live in the platform’s own database; the narrative is written by a tired person on a deadline. A local model with a read-only connection to that database can draft the quarterly narrative from the numbers themselves, flag where this quarter contradicts last, and produce the Spanish and English versions together. Donor records, participant identities and financials never leave the organization’s own machine to do it.

donor data never crosses the internet
↳ and the organizations we have not met yet, where the argument is even shorter:

Anyone whose data is someone else’s secret.

A legal-aid clinic. Privilege does not survive a paste into a consumer chatbot, and the bar does not care that the training setting was off. A model that is not connected to the internet can still summarise intakes, review documents and draft motions.

A community health center. Running locally does not make anyone HIPAA compliant. What it does is remove one business associate, and in practice that is often the one nobody had a signed agreement with. Visit notes get summarised, prior-auth letters get drafted, and Spanish patient instructions get written at a sixth-grade reading level, all of it inside the clinic.

A cooperativa or a credit union. Member financial records, loan files and collections notes. The model reads the file and drafts the letter, and the compliance officer can point at the machine it ran on.

A school. Student records, IEPs, behaviour notes and the endless parent communication, drafted in both languages. Because the child’s file never leaves the principal’s office, there is no consent form to send home about it.

A newsroom. Source material, unpublished drafts, leaked documents. The model that helps you read a thousand pages of a public-records dump should not be run by a company that can be subpoenaed for the prompt.

A small manufacturer, a design office, an accounting firm in March. Contracts, drawings, client books. The pattern is the same every time: the work is valuable because it is confidential, and confidential things should not be typed into someone else’s computer.

The box stays the same.
The models get better.

↳ why the box gets better after you have bought it

The strongest argument for owning the machine is the direction things have been moving for three years: the same quality of answer keeps needing less computer. A model that needed a rack of data-center cards in 2024 has an equivalent that runs on a desk in 2026, and nothing about how these models are built suggests that stops. Three things are driving it, and each one arrives on a machine you already own as a free download.

Fewer parts awake at onceSparse models

The newer architectures only wake a small fraction of the model for each word they write. A model with hundreds of billions of parameters can answer using a tenth of them at a time, which is how something that large runs at conversational speed without a data center. Two years ago the equivalent model used everything, every word, and crawled.

Fewer bits per numberCompression

Weights are stored in fewer bits than they were trained in. The loss is measurable and, for the work on this page, usually not noticeable, and the technique keeps improving. The practical effect is that the same memory holds a bigger, better model every year.

Small models taught by large onesDistillation

The laboratories now use their largest models to train their smallest, and the smallest inherit far more of the capability than anyone expected. The model that fits comfortably on a desk today is roughly what the flagship was a year and a half ago, and the interval keeps shortening.

↳ what that means for a machine bought this year:

A better model every few months, on the same machine.

A cloud plan is the same price next year and the year after, and rises with headcount. A box in your closet costs what it cost, once, and every few months a better model appears that runs on it for nothing. The machine you buy for the work above will, on the same memory and the same electricity, run something noticeably stronger by the time you have finished rolling it out. What the money buys is the memory, and every model that will fit in it for as long as the box runs.

The limit is that the frontier moves too. The very best closed model will stay ahead of the best open one, probably always, and the gap on the hardest problems will not close to zero. What closes is the gap on your problems, because the line for “good enough for this work” keeps moving down in size while your work stays the same size. The two-tier rule below is therefore a permanent part of the design: the local model handles more each year, and the outbound lane gets quieter.

“beholden to any one solution”
Andy Markus, AT&T’s chief data and AI officer, on what the move to open models is meant to avoid, to the Wall Street Journal, August 2026

A company of that size is moving in the direction this page describes, for the reasons this page gives.

↳ and it is not only small organizations moving this way:

AT&T is doing it at telco scale.

In August 2026 the Wall Street Journal reported that open-weight models already handle about a quarter of AT&T’s AI usage, and that its chief data and AI officer expects them to reach seventy to eighty percent. Earlier in the year he told PYMNTS the company had cut its AI costs by ninety percent and tripled its throughput by moving work from large closed models to small ones it runs itself, and that on a given domain it finds a small model “just about as accurate, if not as accurate” as a large one. The reasons he gives are cost and control: the models are theirs to run where they choose, on their own data, without depending on any one vendor’s pricing, terms or roadmap.

The logic that holds for a company of AT&T’s size holds harder for a clinic or a foundation, which has less leverage over a vendor and more to lose from a single leak. We think this becomes the norm: the organization owns the model, the model comes to the data, and the cloud handles only the questions that have had the records taken out of them first. AT&T got there for cost and control. The organizations on this page have those two reasons and one more: the people whose records they hold never agreed to have them anywhere else.

Sources: PYMNTS, 27 February 2026, interview with Andy Markus; the Wall Street Journal’s 17 August 2026 report by Belle Lin, as summarised by The Fly. The usage shares are AT&T’s own figures, not audited.

How it
would work.

↳ three parts, and the third one is the one we would spend the time on

The machine and the model are the easy parts; they are purchases. What makes the thing usable by a team, and safe to leave running, is the layer between them and the people, and that layer is the design job. We make the same argument about the conversational CMS: generation is cheap now, and most of the work is in the rails around it.

Part oneThe box, in the closet

The Mac Studio, on the office network, with the disk encrypted, automatic updates on, a backup that somebody has tested, and the physical lock the closet already has. It runs the model as a service that other machines on the network can talk to. It does not accept connections from the internet, at all, and we would put that in the contract.

Part twoThe connectors

The model is only useful if it can read your things. A shared folder, the case-management database, the archive’s transcripts, the platform’s reporting tables. Each is connected through a standard socket (the same MCP protocol we describe on the CMS page), read-only unless there is a reason, and scoped to what each role is allowed to see.

Part threeThe interface, and the rules

A chat window on the office network, in Spanish and English, with a login that is your existing one rather than a new account with anyone. Behind it, the rules: who can ask about which records, what the model may never write out in full, which answers need a named person’s approval before they become a letter. Writing those rules down, and holding the software to them, is most of the job.

Stays in the room, always

Anything with a name in it.

  • Case files, intake forms, applications, mentor notes
  • Interview recordings and transcripts
  • Donor, member and participant records
  • Financials, contracts, anything under privilege
  • The conversation history itself
Allowed out, by policy, with a log

Questions with the names removed by a person.

  • A hard reasoning problem where the local model is not enough, rewritten without the record
  • Public-source research: what the law says, what a funder published
  • Code and templates that contain no data
  • Nothing automatically. The tiering is a rule a person set, and every outbound question is logged.
↳ the two-tier rule, which is how we would handle the frontier gap:

The honest answer to “is the open model as good as the best closed one” is not on everything, and it does not need to be. The design is a two-tier rule. Everything that touches a record happens on the box. When a question genuinely needs the largest closed model, and some do, the local model helps you write a version of the question with nothing identifying in it, a person reads that version, and only then does it leave. The record never travels. The question sometimes does, stripped, on purpose, with a log. It is a rule your board can read and your auditor can check against the log.

Where it is harder
than the pitch.

↳ six limits that stay true after the box is installed

A page like this one is a pitch whether we want it to be or not, so here is the other column. Everything below is true of the design we just described.

1

Local is not the same as secure

Privacy from the vendor is one property. A box in a closet still has to be patched, backed up, encrypted, physically locked and reachable only by the people who should reach it. If the closet is open and the password is on a sticky note, you have built a very private machine that anyone can walk up to. The security page applies in full.

the closet needs a lock
2

It is one machine

When it is off, the assistant is off. When the disk fails, the models are re-downloadable but the conversations and connectors are only as safe as the backup. It will not stand up as a public-facing product and it does not scale to a thousand users. It suits a team in one building.

a team, not a public
3

It runs at reading speed

The models that suit this work answer at the pace of a person reading aloud, for one or a few people at once. The largest ones are slower than that. Nobody will mistake it for the cloud tool’s instant reply, and for a caseworker drafting a letter that does not matter. For a room of twenty people at once, it does.

reading speed, a few at a time
4

The gap to the frontier is real

On the hardest reasoning the closed models are ahead, and the very largest open ones do not fit on a desk. For the work on this page the gap rarely shows, and it narrows every year. We are not going to tell you it is zero.

months behind, not years, and not zero
5

Read the licence, every time

MIT and Apache are fine. Some models carry custom community licences with usage caps or field-of-use terms, and the licence can change between versions. A model is a file with terms attached, and somebody has to read them before the file goes on the machine.

the licence is the document
6

Somebody has to run it

Models update, runtimes update, connectors break when the database schema changes, and the log needs a reader. This is a fractional-CIO shape of job, a few hours a month, and if nobody owns it the box goes stale the way an unpatched website does, quietly, while everything still looks fine.

a named person, every month
↳ and the part about compliance, which people want to hear and should not:

Running a model locally does not make an organization compliant with anything. HIPAA, FERPA, Puerto Rico’s breach-notification law, a funder’s data clause: each is a set of obligations about how data is handled, and a local model changes one line in the diagram. It removes a third party that was processing the data, which is a large improvement and the one this page is about. It is still not a certificate. Anyone who tells you the box is the compliance is selling the box.

What is on
the invoice.

↳ what the lines are, and what we can and cannot price

We will not put a per-seat number on the cloud alternative, because enterprise AI pricing is negotiated, changes quarterly, and anyone quoting a figure across the board is guessing. What we can do is name the lines. The cloud version is a monthly plan per person, rising with headcount, for as long as you use it, plus the redaction tooling if you are handling sensitive material responsibly, plus the legal review of each vendor’s terms. The local version is a machine, once, and a person, monthly.

~$11k
once, for a desktop machine with the most memory Apple sells.
Roughly the list price this month, before tax. Memory is the only number that matters, because the model has to fit in it whole; we would not economise there. Everything else about the box is ordinary.
$0
per seat, per question, per month, to any vendor.
The models are free to download and free to use commercially under their licences. The electricity is real but small: it is one desktop computer.
1
named person, a few hours a month, to keep it patched, backed up and connected.
This is the line most pitches omit and the one that decides whether the thing is still working in a year. It is the part we would do.
↳ where we come into it, which is narrower than a product:

The machine is Apple’s. The models are published by laboratories that have never heard of you. The runtime is open source. What is specific to your organization is the rest: deciding which records the model may read and which roles may ask, wiring the connectors to the systems you already have, writing the two-tier rule and the log, building the interface your team will use instead of routing around, and running the box afterwards. That is most of the work, it is as much design as engineering, and it is the part that does not come in a box.

We would start, as we always do, by sitting with the team and finding the three questions they most want to ask their own data and cannot. If none of those questions is about a record with a name in it, you do not need this page and a business plan on a cloud tool will do. If all three are, the next step is a trial rather than a purchase order. We rent a machine in a cloud service with the same memory profile as the one we would buy, put the candidate open models on it, and run them against those three questions. The material for that trial is synthetic or de-identified, because a rented machine sits in someone else’s data center and real records do not go there. What comes back tells us which model fits the work and what to expect from it, and only then is it worth putting real numbers against the lines above.

as always ↳

This is an idea we are exploring. The capabilities described are what current open-weight models do on a single desktop machine as of August 2026; the trajectory section is our reading of the last three years, not a guarantee about the next three. Prices are list prices this month and will change. Nothing here is legal advice, and no partner named above has been asked to do any of it. When we have run this for a real team, this page will say what we found.

↳ only if you hold records you cannot paste anywhere

If the paste is the problem.

Most organizations should use a good cloud tool with the training setting off. If yours holds the kind of file this page is about, and your team has stopped asking questions of it because they could not, that is the conversation to have before any hardware is bought.