Support & Education For whoever signs the data agreement An idea we are exploring Field notes · August 2026

Your own model,
on your own desk.
Nothing leaves the room.

Every useful thing an organization does with AI today begins the same way: somebody pastes something into a computer that belongs to someone else. A case note, an intake form, a grant report, an interview transcript, a spreadsheet of families. The tool is excellent, the answer comes back in seconds, and the thing you pasted is now on a server you do not control, under terms you did not read, in a country you did not choose. For a marketing team that is a policy question. For a foundation holding the case files of four hundred families, or the recorded voices of people who trusted an archive with their memories, it is the whole question.

There is a version of this where the model is a file, the file lives on one machine in your office, and nothing you type into it goes anywhere. Not to a vendor, not to a data center, not through a redaction step somebody has to get right every time. The models are open weights anyone can download. The machine is a desktop computer. These are notes on what that adds for the organizations we work with, what it can do today, why it gets better on its own, and the parts of it that are harder than the pitch.

Skip to what we would build with it ↓
Read this as a proposal.
The capabilities are what we can switch on now. The trajectory is our read of three years of watching it. When we have run this for a real team, this page gets their numbers.

What “private”
actually means here.

↳ the word gets used for three different things. this is the strict one.

Every AI vendor will tell you your data is private. What they mean varies. The consumer tier means it trains on what you paste unless you find the setting. The business tier means it does not train, and the retention is whatever the contract says, and the data still crosses the internet and sits on their machines while it is processed. Both of those are real forms of privacy and we have a whole page on choosing between them. This page is about the third thing, which is a different category rather than a better tier.

A model is, physically, a very large file of numbers. Running it means loading that file into memory and doing arithmetic on it. If the file is on a machine in your office and the arithmetic happens there, then the sentence you typed and the document you attached were never transmitted to anyone. There is no vendor in the diagram. There is no retention policy, because there is no one to retain it. There is no terms-of-service update to watch for. The privacy is not a promise somebody made you. It is a property of where the cable goes.

Where the data goesNowhere

The question and the document are read from your disk into your memory, the answer is written back to your screen. The machine can have its network cable pulled and the assistant keeps working, which is a test you can run yourself and which no cloud tool can pass.

What the vendor seesThere is no vendor

The people who trained the model published the weights and walked away. They do not know you downloaded it, they do not know what you ask it, and they have no way to find out. That is what open weights means, and it is why the licence, not the vendor relationship, is the document to read.

What it costs to askElectricity

There is no per-question charge, no per-seat charge, no monthly plan. The cost of a question is the power the box draws while it thinks. That changes behaviour in ways worth noticing: nobody rations questions, and nobody quietly uses the free consumer tool instead because the paid one ran out.

↳ the workaround this replaces, and why it was never good enough:

You stop scrubbing the data before you use it.

The standard advice for using a cloud model on sensitive material is to strip the sensitive part first. Replace names with placeholders, pseudonymise the identifiers, tokenise the account numbers, and send the model the hollowed-out version. An entire product category exists to do this, and organizations pay for it. It has two problems that never go away.

The first is that it has to be right every single time, and it is done by a rule or a small model that does not know that Doña Carmen from the second floor is a name. One miss and the record you were protecting is on the server anyway, and you do not know which one. The second is quieter and worse: the context you scrubbed out was the reason the question was worth asking. A navigator wants to know which of her families is at risk of a benefit cliff next month. Anonymise the families and there is no question left.

With the model in the room, you do not scrub. You ask the real question about the real record, in the words you would use to a colleague, and the record stays where it was.

The data does not go to the model. The model comes to the data.

What it can do
for you, today.

↳ not a roadmap. things we can switch on this quarter.

The reason this is a page now and not a prediction is that the open models crossed a line in the last year. The weights the laboratories publish for free are, for the everyday work of an organization, as good as the paid tools were eighteen months ago, and they run on a desktop computer that plugs into a wall. The hardware is a purchase, not a project: one machine with as much memory as Apple will sell, in the closet where the router lives. What matters is what a team can actually do with it, so here is that list, with an honest note on how well each one works.

1

Read your documents and answer from them

Point it at a folder, a database or an archive and ask questions in plain Spanish or English. It answers from your material and cites where it found it, which is the difference between an assistant and a guess.

works well. the citation is the part to insist on
2

Summarise, draft and translate, in both languages

A case note into a follow-up letter. A quarter of figures into a funder narrative. A policy memo into the Spanish a family will actually read. Register and tone are set once, by you, and held.

works well. review before it leaves the building
3

Pull fields out of forms, scans and photos

The current models see. An intake form photographed on a phone, a scanned letter from 1962, a family’s photograph with a note on the back: names, dates, places and amounts come out as structured fields a person confirms.

good, not perfect. a person confirms
4

Transcribe and tag recordings

Hours of interviews, meetings or voice notes become searchable text with speakers and timestamps, tagged by topic, without the audio going to a transcription service. Puerto Rican Spanish included.

works well. dialect and noise need a check
5

Sort, triage and flag

Applications by readiness, tickets by urgency, the participant who has gone quiet, the record that contradicts last quarter. It reads everything so that a person reads the right things.

works well. the rule is yours, not the model’s
6

Do small pieces of work end to end

Given a connector and a rule, it can draft the letter, file it in the right folder and put the reminder on the calendar, then stop for approval. Longer chains of steps are where it still needs a person alongside.

early. short chains yes, long ones not yet
Real, and good enough

For the work above

The open models running on one desk today do all six of those at a level that would have been the best available in the world eighteen months ago. For a case file, a grant report or an archive, the gap to the most expensive cloud model rarely matters, and it never matters more than the record staying home.

Real, unfixed

For the hardest reasoning

The largest closed models are still ahead on the hardest problems and the very largest open ones do not fit on a desk. If a task genuinely needs that tier, the question goes out and the record stays home. We describe how below, and we would rather design for the gap than deny it.

What we would build
with it, by partner.

↳ dreamed up, not promised. none of these organizations has signed anything.

This is the part of the page we most wanted to write. The organizations we build for share a shape: small teams, a great deal of trust placed in them by the people they serve, and files those people would not want in a stranger’s data center. Every one of them has a use for a model. Almost none of them can responsibly paste their real work into one. Here is what changes when the model is in the room, told through the work we already know.

1

Family Navigator: the caseload that can finally be asked a question

Vimenti’s navigators at the Boys & Girls Clubs carry families through benefits, housing, schooling and work, and the record of each family is the most sensitive thing in the building. Today a navigator who wants to know which of my families has a recertification due before October, and which of those also lost hours last month has to know it already, or build a spreadsheet. With the model reading the case files on a machine in the club, that is a sentence typed in Spanish on a Tuesday. Drafting the follow-up letter in the register the family expects is the next sentence. The files never leave the club, and the navigator never opens a browser tab to a vendor.

case files stay in the building
2

The benefit-cliff calculator, explained to the family it applies to

The Calculadora de Beneficios, built with the Instituto del Desarrollo de la Juventud and PSI, computes what a raise costs a family in lost assistance. The arithmetic is the easy part. The hard part is explaining, to a specific mother with a specific job offer, what the number means and what her options are, in plain Puerto Rican Spanish, without a caseworker in the room. A local model that has read the rules and her actual figures can do that explanation, and can draft the testimony a policy analyst needs for a hearing next week from a hundred real cases, none of which had to be anonymised first.

real cases, not de-identified ones
3

Fundación Rimas: reading applications from minors without a vendor in the loop

La Academia’s participation platform collects applications, mentor notes and progress records about young people, and we have written before about what it means to have minors in a dataset. A model that helps staff triage applications, summarise a mentor’s notes before a check-in, or spot a participant who has gone quiet is useful. Sending a fifteen-year-old’s essay to a cloud vendor to get that help is a decision a board should have to make out loud. On a machine in the foundation’s office the decision is not required.

nobody’s child in anyone’s training set
4

La Voz del Centro: asking a question across a thousand hours of voices

The archive holds decades of interviews with people who agreed to be recorded for a purpose. Transcribing them, tagging them and making them searchable is already work we do, and the processing stack is described on that page by what it does rather than who sells it. A model in the room takes the next step: find every time anyone describes the 1985 landslide, in their own words, or what did the interviewees born before 1940 say about the sugar mills, answered from the actual transcripts with citations back to the minute. The voices never go to a transcription vendor, which is what the interviewees were promised.

the voices stay with the archive
5

Archivo Negro: cataloguing family memory without exporting it

Families hand El Archivo Vivo photographs, letters and stories, and the curators contextualise and safeguard them. The bottleneck is description: every item needs a caption, a date estimate, names, places, a link to the collection it belongs in. A vision-capable local model can draft that description from the scan and the note the family sent, for a curator to correct, at the pace the submissions arrive. The family’s photograph of their grandmother is never uploaded to a company that will keep a copy, which is a thing the archive can now say to the family in writing.

the grandmother’s photo stays in Puerto Rico
6

PSI’s front door: the grant report that writes its first draft from the data

A platform for social impact runs on reporting: to funders, to boards, to the public. The numbers live in the platform’s own database; the narrative is written by a tired person on a deadline. A local model with a read-only connection to that database can draft the quarterly narrative from the actual figures, flag where this quarter contradicts last, and produce the Spanish and English versions together. Donor records, participant identities and financials never leave the organization’s own machine to do it.

donor data never crosses the internet
↳ and the organizations we have not met yet, where the argument is even shorter:

Anyone whose data is someone else’s secret.

A legal-aid clinic. Privilege does not survive a paste into a consumer chatbot, and the bar does not care that the setting was off. Intake summaries, document review and draft motions from a model that has never seen the internet.

A community health center. Running locally does not make anyone HIPAA compliant; it removes one business associate from the diagram, which is usually the one nobody had a signed agreement with. Visit notes summarised, prior-auth letters drafted, Spanish patient instructions written at a sixth-grade reading level, in the clinic.

A cooperativa or a credit union. Member financial records, loan files and collections notes, with a model that can read a file and draft the letter, on a box the compliance officer can point to.

A school. Student records, IEPs, behaviour notes and the endless parent communication, drafted in both languages by a model that no parent has to be asked to consent to, because their child’s file stayed in the principal’s office.

A newsroom. Source material, unpublished drafts, leaked documents. The model that helps you read a thousand pages of a public-records dump should not be run by a company that can be subpoenaed for the prompt.

A small manufacturer, a design office, an accounting firm in March. Contracts, drawings, client books. The pattern is the same every time: the work is valuable because it is confidential, and confidential things should not be typed into someone else’s computer.

The box stays the same.
The models get better.

↳ this is the part that makes it a good purchase rather than a clever one

The strongest argument for owning the machine is not what runs on it this year. It is the direction everything has been moving, consistently, for three years: the same quality of answer keeps needing less computer. A model that needed a rack of data-center cards in 2024 has an equivalent that runs on a desk in 2026, and nothing about how these models are built suggests that stops. Three specific things are driving it, and each one lands on a machine you already own as a free download.

Fewer parts awake at onceSparse models

The newer architectures only wake a small fraction of the model for each word they write. A model with hundreds of billions of parameters can answer using a tenth of them at a time, which is how something that large runs at conversational speed without a data center. Two years ago the equivalent model used everything, every word, and crawled.

Fewer bits per numberCompression

Weights are stored in fewer bits than they were trained in. The loss is measurable and, for the work on this page, usually not noticeable, and the technique keeps improving. The practical effect is that the same memory holds a bigger, better model every year.

Small models taught by large onesDistillation

The laboratories now use their largest models to train their smallest, and the smallest inherit far more of the capability than anyone expected. The model that fits comfortably on a desk today is roughly what the flagship was a year and a half ago, and the interval keeps shortening.

↳ what that means for a machine bought this year:

The hardware is fixed. The capability is not.

A cloud plan is the same price next year and the year after, and rises with headcount. A box in your closet costs what it cost, once, and every few months a better model appears that runs on it for nothing. The machine you buy for the work above will, on the same memory and the same electricity, run something noticeably stronger by the time you have finished rolling it out. You are not buying this year’s model. You are buying a seat at every model that fits, for as long as the box runs.

The honest limit is that the frontier moves too. The very best closed model will stay ahead of the best open one, probably always, and the gap on the hardest problems will not close to zero. What closes is the gap on your problems, because the line for “good enough for this work” keeps moving down in size while your work stays the same size. That is why the two-tier rule below is a permanent part of the design rather than a stopgap: the local model handles more each year, and the outbound lane gets quieter.

“beholden to any one solution”
Andy Markus, AT&T’s chief data and AI officer, on what the move to open models is meant to avoid, to the Wall Street Journal, August 2026

That is not a foundation with eleven staff. It is one of the largest companies in the United States, and it is moving in the direction this page describes, for the reasons this page gives.

↳ and it is not only small organizations moving this way:

AT&T is doing it at telco scale.

In August 2026 the Wall Street Journal reported that open-weight models already handle about a quarter of AT&T’s AI usage, and that its chief data and AI officer expects them to reach seventy to eighty percent. Earlier in the year he told PYMNTS the company had cut its AI costs by ninety percent and tripled its throughput by moving work from large closed models to small ones it runs itself, and that on a given domain it finds a small model “just about as accurate, if not as accurate” as a large one. The reasons he gives are cost and control: the models are theirs to run where they choose, on their own data, without depending on any one vendor’s pricing, terms or roadmap.

The logic that holds for a company of AT&T’s size holds harder for a clinic or a foundation, which has less leverage over a vendor and more to lose from a single leak. We think this becomes the norm: the organization owns the model, the model comes to the data, and the cloud is where stripped questions go, not where the records live. AT&T got there for cost and control. The organizations on this page get there for the same two reasons and a third, which is that the people in their files never agreed to be anywhere else.

Sources: PYMNTS, 27 February 2026, interview with Andy Markus; the Wall Street Journal’s 17 August 2026 report by Belle Lin, as summarised by The Fly. The usage shares are AT&T’s own figures, not audited.

How it would
actually work.

↳ three parts, and the third one is the one we would spend the time on

The machine and the model are the easy parts; they are purchases. What makes the thing usable by a team, and safe to leave running, is the layer between them and the people, and that is the design job. This is the same argument we make about the conversational CMS: generation is cheap now, and the scarce thing is the rails.

Part oneThe box, in the closet

The Mac Studio, on the office network, with the disk encrypted, automatic updates on, a backup that somebody has tested, and the physical lock the closet already has. It runs the model as a service that other machines on the network can talk to. It does not accept connections from the internet, at all, and we would put that in the contract.

Part twoThe connectors

The model is only useful if it can read your things. A shared folder, the case-management database, the archive’s transcripts, the platform’s reporting tables. Each is connected through a standard socket (the same MCP protocol we describe on the CMS page), read-only unless there is a reason, and scoped to what each role is allowed to see.

Part threeThe interface, and the rules

A chat window on the office network, in Spanish and English, with a login that is your existing one rather than a new account with anyone. Behind it, the rules: who can ask about which records, what the model may never write out in full, which answers need a named person’s approval before they become a letter. The rules are the product.

Stays in the room, always

Anything with a name in it.

  • Case files, intake forms, applications, mentor notes
  • Interview recordings and transcripts
  • Donor, member and participant records
  • Financials, contracts, anything under privilege
  • The conversation history itself
Allowed out, by policy, with a log

Questions with the names removed by a person.

  • A hard reasoning problem where the local model is not enough, rewritten without the record
  • Public-source research: what the law says, what a funder published
  • Code and templates that contain no data
  • Nothing automatically. The tiering is a rule a person set, and every outbound question is logged.
↳ the two-tier rule, which is where the frontier question is actually settled:

The honest answer to “is the open model as good as the best closed one” is not on everything, and it does not need to be. The design is a two-tier rule. Everything that touches a record happens on the box. When a question genuinely needs the largest closed model, and some do, the local model helps you write a version of the question with nothing identifying in it, a person reads that version, and only then does it leave. The record never travels. The question sometimes does, stripped, on purpose, with a log. That is a rule your board can read and your auditor can check, which is more than can be said for a setting in a vendor’s admin panel.

The parts we are
not going to soften.

↳ read this before the section above starts to sound easy

A page like this one is a pitch whether we want it to be or not, so here is the other column. Everything below is true of the design we just described, and a vendor selling you the same thing would probably leave most of it out.

1

Local is not the same as secure

Privacy from the vendor is one property. A box in a closet still has to be patched, backed up, encrypted, physically locked and reachable only by the people who should reach it. If the closet is open and the password is on a sticky note, you have built a very private machine that anyone can walk up to. The security page applies in full.

the closet needs a lock
2

It is one machine

When it is off, the assistant is off. When the disk fails, the models are re-downloadable but the conversations and connectors are only as safe as the backup. It is not a product you put in front of the public and it does not scale to a thousand users. It is for a team, in a building.

a team, not a public
3

It runs at reading speed

The models that suit this work answer at the pace of a person reading aloud, for one or a few people at once. The largest ones are slower than that. Nobody will mistake it for the cloud tool’s instant reply, and for a caseworker drafting a letter that does not matter. For a room of twenty people at once, it does.

reading speed, a few at a time
4

The gap to the frontier is real

On the hardest reasoning the closed models are ahead, and the very largest open ones do not fit on a desk. For the work on this page the gap rarely shows, and it narrows every year. We are not going to tell you it is zero.

months behind, not years, and not zero
5

Read the licence, every time

MIT and Apache are fine. Some models carry custom community licences with usage caps or field-of-use terms, and the licence can change between versions. A model is a file with terms attached, and the terms are the part that makes it into a contract.

the licence is the document
6

Somebody has to run it

Models update, runtimes update, connectors break when the database schema changes, and the log needs a reader. This is a fractional-CIO shape of job, a few hours a month, and if nobody owns it the box goes stale the way an unpatched website does. Quietly, while looking fine.

a named person, every month
↳ and the part about compliance, which people want to hear and should not:

Running a model locally does not make an organization compliant with anything. HIPAA, FERPA, Puerto Rico’s breach-notification law, a funder’s data clause: each is a set of obligations about how data is handled, and a local model changes one line in the diagram. It removes a third party that was processing the data. That is a genuinely large improvement and it is the one this page is about. It is not a certificate, and anyone who tells you the box is the compliance is selling the box.

What is on
the invoice.

↳ which lines exist, not how big they are

We will not put a per-seat number on the cloud alternative, because enterprise AI pricing is negotiated, changes quarterly, and anyone quoting a figure across the board is guessing. What we can do is name the lines. The cloud version is a monthly plan per person, rising with headcount, for as long as you use it, plus the redaction tooling if you are handling sensitive material responsibly, plus the legal review of each vendor’s terms. The local version is a machine, once, and a person, monthly.

~$11k
once, for a desktop machine with the most memory Apple sells.
Roughly the list price this month, before tax. Memory is the only number that matters, because the model has to fit in it whole; we would not economise there. Everything else about the box is ordinary.
$0
per seat, per question, per month, to any vendor.
The models are free to download and free to use commercially under their licences. The electricity is real but small: it is one desktop computer.
1
named person, a few hours a month, to keep it patched, backed up and connected.
This is the line most pitches omit and the one that decides whether the thing is still working in a year. It is the part we would do.
↳ where we come into it, which is narrower than a product:

The machine is Apple’s. The models are published by laboratories that have never heard of you. The runtime is open source. What is specific to your organization is the rest: deciding which records the model may read and which roles may ask, wiring the connectors to the systems you actually have, writing the two-tier rule and the log, building the interface your team will use instead of routing around, and running the box afterwards. That is a design job with an engineering shape, it is most of the work, and it is the part a vendor cannot sell you in a box.

We would start, as we always do, by sitting with the team and finding the three questions they most want to ask their own data and cannot. If none of those questions is about a record with a name in it, you do not need this page and a business plan on a cloud tool will do. If all three are, the machine pays for itself the first time someone does not have to build the spreadsheet.

as always ↳

This is an idea we are exploring, not a product page. The capabilities described are what current open-weight models do on a single desktop machine as of August 2026; the trajectory section is our reading of the last three years, not a guarantee about the next three. Prices are list prices this month and will change. Nothing here is legal advice, and no partner named above has been asked to do any of it. When we have run this for a real team, this page will say what we found.

↳ only if you hold records you cannot paste anywhere

If the paste is the problem.

Most organizations should use a good cloud tool with the training setting off. If yours holds the kind of file this page is about, and your team has stopped asking questions of it because they could not, that is the conversation to have before any hardware is bought.