La memoria de un país, en su propia voz. The memory of a country, in its own voice: something you can ask, walk and hear, and now something its own author corrects by talking to it.
La Voz del Centro is the Caribbean's first podcast. For more than twenty years, its creator and host Ángel Collado-Schwarz has sat down each week with historians, artists and protagonists to talk through the history and culture of Puerto Rico and the Caribbean. Past 1,100 episodes, it became the most comprehensive oral-history archive of the island that exists, from the Taínos to today. Mutiny Labs has been the technology behind it from day one, in the background. It lived as audio: profound, but linear, and impossible to search by what was said.
Twenty years of recorded conversation is a national memory locked in audio. You could not ask it a question, jump to the minute where something was said, or take in the shape of two decades at a glance. The Fundación Voz del Centro wanted more than a better player: an archive people could interrogate, walk and inhabit, faithful to every source, and unmistakably Puerto Rican rather than a generic media library dressed in someone else's aesthetic.
The archive is not a build that shipped once; it is a system that keeps itself alive. Eighteen automated stages run every week, taking each new episode from the public feed to the finished experience: transcription with millisecond timestamps, annotation against a strict schema, two machine judges that score every chapter and quote before anything publishes, audio clips cut to the exact second, passages indexed by meaning, scenes written and posters generated under the house art direction, English translation that never touches the Spanish source, and a health check that audits every run against the last. Seven AI models are orchestrated across the flow, each chosen for one job. Every stage is idempotent and resumable; nothing partial reaches the public, and nothing is asserted that is not recorded. And the final stage is human: a conversational editor, held by the same rails, through which the host corrects the archive himself.
The archive's owner corrects it by conversation. The host describes a transcription error in his own words; a governed assistant searches the archive and returns a proposal card with before and after, the count per area and the exact minute. Nothing changes until he presses Apply, every cell touched is logged with its previous value and can be undone, and anything beyond a text correction becomes a drafted request to the studio. No dashboard, no fields, no way to break the site.
A floating conversation with the whole archive, in the courteous voice of the program. Ask in plain Spanish and the answer streams in, grounded in the corpus; every claim carries a tap-to-play citation that plays the exact minute of the exact episode without leaving the window. What the archive does not document, it will not say.
Every episode plays with AI-generated chapters marked on the scrubber, a timestamped outline, and a Cartel de Pueblo poster for each chapter that changes as you listen.
The most memorable spoken moments surfaced as full-bleed quote cards over original art, each one a single tap from the exact second in the audio.
Curated narrative routes string episodes into a larger story, stop by stop; a timeline reorders the whole archive by the era each episode narrates, from the Taínos to the 21st century; and three doors open it by theme, by voice and by entity, with people, places, institutions, events and works ranked by how often the archive returns to them.
We parse the public podcast feed into a clean, idempotent manifest keyed to stable IDs. The episode number printed in each title is treated as the authority, because the feed's own numbering field carries typos going back years.
Each episode is mirrored to global object storage with its artwork and verified against the feed's own duration, so the archive owns its masters instead of renting them from someone else's CDN.
Every episode is transcribed in Spanish at millisecond resolution, with two independent speech providers held in reserve behind the primary. Speaker attribution is settled by comparing voices against the host in that same episode, never by guessing from the text.
The recurring introduction, the sponsor read and the sign off are detected automatically and excluded from search, indexing, clips and players. The host's own closing recap was classified as content and deliberately kept. That distinction covers 4,169 spans.
Generative AI reads each transcript under a strict JSON schema: summaries, timestamped chapters, named entities and aliases, the historical period covered, and quotable moments. Timestamps must come from the transcript segments. Inventing one is forbidden.
A mechanical judge scores coverage, richness and anchoring, on purpose not a model grading another model's prose. Then every quote is checked word for word against the transcript and realigned to where the words actually are: 4,604 of 6,751 survive, and 993 clear the bar for the public wall.
Every ninety second passage is embedded into a vector index and fused with Spanish full text search, so a question returns the episode and the precise minute rather than a list of files. The padding never enters the index.
A model writes one visual metaphor per chapter, then generates it under a single art direction, Cartel de Pueblo, in the language of Puerto Rico's mid century public education posters. Scenes hold places, objects and anonymous figures. No likeness of a named private individual is ever attempted.
The whole archive is translated into English beside the Spanish, never over it, with proper nouns left unanglicized. Only then does the publish gate open, and only if transcription and annotation both exist. A failed overnight run cannot leave a hollow page in public.
The last stage is a person. In a private editor, the host talks to an assistant that can only search and propose. Each proposal writes the transcript and the search index together and logs every cell with its previous value; he applies with one button and can undo with another.
The annotating model is required to take its timestamps from the transcript segments, and forbidden to invent one. Every claim the archive makes lands on the exact point in the recording, so a listener can always go and check.
Models paraphrase when you ask them to quote. An independent pass confirms the first and last words appear verbatim inside the window, then realigns the clip to where the words actually are.
Attribution is decided by voice, comparing the moment against the host in that same episode. When an episode has several guests and certainty is not there, the quote goes out unnamed. No name beats a wrong name.
The conversational search answers only from retrieved passages, cites every claim, and is barred from using its own general knowledge even when it knows the answer. Links and audio come from the database, never from model output.
Translation is written into separate columns and never overwrites the source. Proper nouns are preserved without anglicizing and transcription markers are kept literally.
Each episode stores which system produced each layer of its analysis and its translation, so any result can be audited, traced and redone rather than taken on faith.
“La genialidad creativa y tecnológica de Gino Sanchez ha logrado maximizar el recurso de inteligencia artificial para comunicar de una forma dramáticamente accesible y amigable a todo el mundo, la historia, vida y sociedad de la nación puertorriqueña contextualizada dentro de su situación colonial y la realidad global. Al igual que el proyecto de la Voz del Centro, son únicos en el mundo.”
“Archivo histórico con inteligencia artificial” · El Nuevo Día
A full page in Flash & Cultura, by Damaris Hernández Mercado, 31 August 2026.
read the article ↗The conversational CMS
The idea behind El Editor, written up: how a team changes its own site through a governed conversation instead of a dashboard, and the four rails that make it safe to hand to anyone.
read the essay ↗under the hood: Semantic search over managed Postgres with vector indexing, retrieval-grounded conversational AI with streamed, citation-locked answers, AI transcription and speaker diarization, generative-AI enrichment against a strict schema, global object storage with edge delivery, React front end, generative art under one art direction
We reply personally. A senior partner, not a sales pod. We'll tell you straight whether we can help.