---
title: "How We Built an AI Podcast That Writes, Voices and Publishes Itself Every Morning"
description: "How to make an AI podcast that writes, voices, checks, and publishes itself daily with Cloudflare Workers, GitHub Actions, Gemini, and Transistor."
url: https://ziplyne.agency/blog/how-we-built-an-ai-podcast-that-writes-voices-and-publishes
published: 2026-09-28T17:54:31.995Z
updated: 2026-09-28T17:54:31.995Z
author: "ZipLyne Editorial"
publisher: ZipLyne
---

# How We Built an AI Podcast That Writes, Voices and Publishes Itself Every Morning

> How to make an AI podcast that writes, voices, checks, and publishes itself daily with Cloudflare Workers, GitHub Actions, Gemini, and Transistor.

## Key takeaways

- A 05:15 Cloudflare Workers cron kicks off the stack behind this AI podcast: GitHub Actions triggers the build, a scored story list feeds the script, Claude via OpenRouter writes it, Gemini's text-to-speech voices it, ElevenLabs checks pronunciation, ffmpeg mixes it, and Transistor publishes it, live by 05:30.
- Fifteen minutes separate the trigger firing from the episode going live, and hitting that window meant skipping GitHub's own scheduler entirely.
- Story selection here comes from a ranked, time-boxed list of what actually happened, not from letting the model invent or guess at news.
- Claude Opus 5.5, called through OpenRouter, writes this script to a 950-to-1,150-word target for a five-to-seven-minute episode at roughly 170 words a minute.

## What is an AI podcast that publishes itself every morning?

A 05:15 Cloudflare Workers cron kicks off the stack behind this AI podcast: GitHub Actions triggers the build, a scored story list feeds the script, Claude via OpenRouter writes it, Gemini's text-to-speech voices it, ElevenLabs checks pronunciation, ffmpeg mixes it, and Transistor publishes it, live by 05:30.

An AI podcast, in the plain sense, is a show where AI writes the script, voices every line, and pushes the finished episode live with no human editing pass. ZipLyne built one for baba News, an English-language news product about Israel: Good Morning, Israel, a daily two-host show with two named hosts, Maya and Daniel, and a fixed daily structure. Nobody writes the script by hand, records the voices, or uploads the file.

That's the gap between what most guides describe and what actually ships. A single text-to-speech tool reading a script is a demo. A system that picks its own stories, checks its own pronunciation, catches its own failures, and still publishes on a fixed schedule with nobody watching is a production pipeline. **The whole thing runs without a single person touching a keyboard on any given morning.** The rest of this piece walks through how, in the order the pieces actually run, including what broke along the way.

![Diagram: How We Built an AI Podcast That Writes, Voices and Publishes Itself Every Morning](https://pub-aee74429e0604b2da81fc4f8bd15a430.r2.dev/ziplyne-agency/how-we-built-an-ai-podcast-that-writes-voices-and-publishes-79c6883984a134b1.webp)

## How do you automate an AI podcast so it runs without anyone touching a computer?

Fifteen minutes separate the trigger firing from the episode going live, and hitting that window meant skipping GitHub's own scheduler entirely. The build fires from a Cloudflare Workers cron job instead, because GitHub's native scheduled runs arrived hours late in testing.

Here is the actual chain, in order:

1. **A Cloudflare Workers cron fires at 05:15 Israel time.** The Worker checks Israel's local hour internally, so daylight saving never shifts the trigger even though the underlying cron clock runs in UTC.
2. **The Worker calls GitHub's workflow\_dispatch API**, using a token scoped to start workflows in one repository only, and kicks off the build on GitHub-hosted Linux runners. No office machine has to be powered on, connected, or awake.
3. **GitHub's own scheduled trigger stays wired in as a backup, not the primary path.** The real trigger moved to Cloudflare specifically because GitHub Actions' native `schedule` events proved unreliable for a job that has to run at a precise minute every morning.
4. **The pipeline runs through story selection, scripting, voicing, mixing, artwork, and publishing**, and the episode is live by 05:30.

📊

By the numbers

The gap between trigger and live episode is fifteen minutes, from a 05:15 Israel-time cron fire to a published episode by 05:30.

The lesson generalizes past podcasts: any daily AI system that has to run unattended needs a trigger that doesn't depend on a platform's best-effort scheduler, and a token scoped tight enough that a compromised key can't touch anything outside the one job it's meant to run.

## How does an AI podcast pick its stories and avoid repeating yesterday's news?

Story selection here comes from a ranked, time-boxed list of what actually happened, not from letting the model invent or guess at news. Good Morning, Israel pulls from the same scored story list that drives the baba News homepage: hard news published within the last 26 hours, ranked by importance with a recency decay, cut to the top 8 stories.

That single design choice does two jobs at once. First, an editor pinning or suppressing a story on the website automatically changes what the show covers that morning, with no separate editorial step for audio. Second, the 26-hour window keeps the episode from ever leading with something listeners already heard two days ago.

Continuity is the harder problem, and it's the one most AI content pipelines skip. Before writing a word of new script, the model reads the full transcripts of the last three episodes, pulled straight from the podcast host's public transcript pages. That gives it three concrete jobs:

1. Recognize a story it already covered and treat it as a continuation, not breaking news.
2. Say explicitly what changed since yesterday, rather than restating the same facts as if they were new.
3. Keep the sign-off and callback lines consistent with what listeners heard the day before.

Most AI news pipelines generate each episode in isolation. Reading three days of its own transcripts before writing is what keeps this show from sounding like it forgot what it said yesterday morning.

## How do you write a script for a two-host AI podcast?

Claude Opus 5.5, called through OpenRouter, writes this script to a 950-to-1,150-word target for a five-to-seven-minute episode at roughly 170 words a minute. That number matters more than the model choice: a fixed word count and a fixed structure are what keep a two-host script from wandering.

The two hosts have fixed jobs. Maya anchors, delivering the top-line story. Daniel explains the background and context underneath it. That division stays constant every day, which is part of what makes the show sound like a real program instead of a randomly assembled monologue split in two.

The daily structure never changes either:

1. Cold open on the single biggest story of the morning.
2. Greeting and show intro.
3. Top stories covered in depth.
4. A quick round-up of smaller items.
5. A sign-off naming one specific thing to watch.
6. A short call to follow the show.
7. "See you tomorrow morning."

Writing for a voice model instead of a reader means specific rules: numbers, dates, and money get spelled out in words instead of symbols, acronyms are written letter by letter with spaces, sentences stay short, and commas land where a human speaker would actually breathe.

Before anything gets voiced, the script, episode title, and show notes all run through a regular expression checking against the show's banned-vocabulary list. A script that trips it gets rewritten once. If it fails a second time, it's refused outright rather than published anyway. That's the editorial gate that lets the rest of the pipeline run unsupervised: catch the problem in text, before it ever becomes audio.

Ready to see how ZipLyne can help?

[Let's build something real](https://ziplyne.agency/contact)

## How do you make AI podcast voices sound natural and consistent?

Natural, consistent voices come from locking two things down: the same voice IDs every episode, and per-line direction instead of one generic style prompt for the whole script. The show runs on Google's Gemini text-to-speech in multi-speaker mode, built specifically for exact text recitation with fine-grained style control, which Google positions directly for use cases like podcast generation.

Two fixed prebuilt voices carry the whole show: Kore and Charon. Each line of dialogue gets sent as its own tagged part, attributed to whichever host is speaking. Google's April 2026 prompting guidance for Gemini 3.1 Flash TTS documents more than 200 audio tags for controlling pace, tone, and expression, plus 30 prebuilt voices across more than 70 languages, and that's the toolkit this pipeline draws from.

There's a real wrinkle worth knowing before building one of these: the model has no separate channel for stage direction. It accepts no distinct "notes" block. So director's-note style instructions, a specific pacing or a particular accent, have to travel inside each line's own style field. Testing showed the model follows those notes rather than reading them aloud, but only when they're attached at the line level. Sparse tags like \[chuckles\] or \[serious\] add texture, but only away from serious news; a cheerful tag on a grim story breaks the illusion instantly.

The bigger fix came from a failure. Sending the whole script as one voicing request let the two voices drift over the course of an episode, and in one run a listener heard what sounded like a third voice appear. The fix was batching: lines are voiced 12 at a time, with each batch retried on failure and split in half if a content filter blocks it, then joined as raw audio under a single WAV header. Batching in chunks that small is what keeps Kore sounding like Kore for the entire five to seven minutes.

## Keep reading

- [What Real AI App Development Looks Like in Production](https://ziplyne.agency/blog/what-real-ai-app-development-looks-like-in-production)
- [Why Your Agency Should Automate Client Blog Publishing in](https://ziplyne.agency/blog/why-your-agency-should-automate-client-blog-publishing-in-2026)
- [How to Choose the First AI Workflow to Build: A Scorecard](https://ziplyne.agency/blog/how-to-choose-the-first-ai-workflow-to-build-a-scorecard)
- [How to Turn a Repetitive Process Into a Custom AI System](https://ziplyne.agency/blog/how-to-turn-a-repetitive-process-into-a-custom-ai-system)

## How do you check an AI podcast for pronunciation mistakes before publishing?

Four takes in a row: that's how many tries it once took the voice model to say the show's own brand name correctly, rendering it at one point as "Habada News." The fix isn't a better voice model. It's a second AI listening to the first one before anything gets published.

The listening-check loop runs on every batch of voiced audio:

1. ElevenLabs Scribe transcribes the batch with word-level timing data.
2. The system checks that the brand name was heard correctly, as a whole word, not spelled out letter by letter.
3. It checks that the spoken brand mention lasts under 0.8 seconds.
4. It checks that at least 85 percent of the intended words are present in the transcript.
5. Any batch that fails gets re-voiced, up to five takes total.
6. If it still fails after five takes, the episode does not publish.

ElevenLabs' Scribe v2 model supports precise word-level timestamps as a core capability, which is exactly what makes a check like "did the brand name land in under 0.8 seconds" possible to automate at all.

⚠️

Watch out

A general multimodal model was tried first as the listener, and it gave opposite answers on the same audio clip on repeated runs. Swapping in a dedicated speech-to-text model made the check consistent. If you're building a similar QA loop, don't assume any model that can "hear" audio is a reliable judge of it.

Because the brand line is the one sentence that can never be wrong, it isn't generated fresh every morning at all. It's recorded once, verified, and spliced into every episode, the same way a radio station records its station ID once and reuses it forever. Everything else in the episode, including the date, gets voiced new each day. Deciding what absolutely cannot be wrong and verifying it directly, instead of hoping the model got it right, is the same discipline behind [AI systems that return checkable values instead of prose](https://ziplyne.agency/blog/ai-that-doesnt-talk-typesafe-jev-guide).

## How do you mix, illustrate, and publish an AI podcast automatically?

Automatic mixing and publishing means every step after the voice audio exists still runs without a human, from the music bed to the final upload. The audio gets assembled with ffmpeg: a theme opens each episode with about five seconds of music, a soft bed ducks under the first spoken lines, the theme returns to close the show, and the whole mix is loudness-normalized to minus 16 LUFS before exporting at 192 kbps.

Artwork gets the same automated treatment, in two steps:

1. Claude Sonnet 5 writes a short art brief pulled from the day's top stories, one focal object for the lead story plus one or two supporting details for the next stories down, following fixed rules: no people, no text, no flags, no religious symbols, nothing violent.
2. OpenAI's GPT Image 2.5 Sunburst renders that brief into a 2048-by-2048 image through fal. If the render fails for any reason, the episode falls back to the show's standard cover and still publishes on schedule.

Publishing runs through the Transistor API: the finished audio uploads, the episode gets created with its full transcript, episode number, keywords, and show notes that link each story to its own page, and the artwork attaches before the episode is published. Transistor's own distribution then pushes the feed out to Apple Podcasts, Spotify, and other listening apps, and the same RSS feed can be picked up by video platforms too. Transistor's API documentation notes a rate limit of 10 requests per 10 seconds, which matters for a scripted daily batch upload rather than a dashboard click-through.

This is the part most guides skip entirely: the tool stack matters less than the fact that mixing, artwork, and publishing all have to succeed, or fail gracefully, with nobody standing by to notice. For a broader look at the model stack behind builds like this one, see this roundup of [AI tools built for real production work](https://ziplyne.agency/blog/ai-tools-10x-business-productivity-2026).

## What breaks when you automate a podcast, and how do you fix it?

Every automated podcast pipeline breaks in the same handful of predictable places: pronunciation, voice consistency, scheduling, dead links, dead air, and image cropping. Good Morning, Israel hit all six failure modes during development, including a brand name mispronounced four takes in a row and a scheduled trigger that arrived hours late. Each failure produced a specific, permanent fix rather than a one-off patch, and the table below lists all six alongside what changed.

| Failure                | What happened                                                                              | Fix that stuck                                                            |
| ---------------------- | ------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------- |
| Brand mispronunciation | Voice model said "Habada News," wrong four takes in a row in one test                      | Brand line recorded once, verified, spliced into every episode            |
| Voice drift            | A full-script voicing request produced a perceived third voice mid-episode                 | Voicing batched to 12 lines at a time, joined as raw audio                |
| Late scheduled runs    | GitHub's native scheduled trigger started hours late in testing                            | Moved the trigger to a Cloudflare Workers cron calling workflow\_dispatch |
| Dead link              | A "read today's stories" link pointed at a page not written until an hour after publishing | Replaced with direct links to each story's own page in the show notes     |
| Dead air               | Roughly five seconds of silence sat between the intro music and the first spoken word      | Voice timing synced directly to the music cue                             |
| Cropped artwork        | Landscape story images got badly cropped when forced into square podcast art               | Replaced with artwork generated square from the start                     |

The pipeline also runs on a small set of reliability rules that assume something will eventually go wrong on any given morning:

1. **One episode per day, checked first.** The system checks whether today's episode already exists before running, so two triggers firing close together can't produce duplicates.
2. **Publish the half-finished draft rather than rebuild.** If a run fails partway through, the system publishes whatever draft state it reached instead of starting the whole pipeline over.
3. **A failure opens a GitHub issue automatically.** Nobody has to notice a broken run by listening to a missing episode.
4. **Too few qualifying stories stops the run with a clear alert**, rather than publishing a thin, padded episode just to hit the schedule.

None of these fixes came from a spec document. They came from the show actually breaking in production and someone deciding the failure could never happen the same way twice.

## How do you build a self-publishing AI podcast for your own business?

Eight steps separate a real source of truth from a published episode, and skipping the first one is why most attempts stall. A product changelog, CRM notes, an industry news feed, or your own CMS all work the same way baba News's scored story list works for Good Morning, Israel: they give the script model something real to write about instead of something to invent.

From there, the build order looks like this:

1. **Pick the source of truth first.** Your CMS, changelog, CRM activity, or an industry feed, whichever one actually reflects what happened this week.
2. **Fix the format and house rules before writing any script.** Decide the structure, the word target, and the banned-vocabulary list up front, the same way Good Morning, Israel locked a seven-part daily structure and a 950-to-1,150-word target.
3. **Write for the ear, not the page.** Numbers spelled out, acronyms spaced letter by letter, short sentences, commas where a breath goes.
4. **Choose voices and lock them.** Pick fixed voice IDs and keep them, and if two hosts, write different jobs for each one rather than splitting one voice's lines in half.
5. **Add an automatic listening check.** Have a second, speech-specific model transcribe the audio and flag anything that doesn't match the intended words before it ships.
6. **Record the one line that must never be wrong.** A brand name, a legal disclaimer, anything a wrong pronunciation would embarrass you for, splice it in instead of generating it fresh.
7. **Run it on a scheduler that doesn't depend on a laptop.** A cloud cron calling a workflow API, not a cron job on someone's desk machine.
8. **Publish through a podcast host with a real API.** Upload, attach metadata and artwork, and let the host's own distribution handle the podcast apps.

This is the same discipline behind any production AI system, not just a podcast: pick the real source of truth, lock the rules, check the output, and let it run without anyone babysitting it. The same pattern shows up in builds that [replace 20 hours a week of manual admin work](https://ziplyne.agency/blog/built-ai-system-replaced-20-hours-weekly-admin-work) instead of a daily show, proof that the discipline transfers across use cases.

## Frequently asked questions

### What tools do you need to build an AI podcast like this one?

Building a fully automated AI podcast takes six or seven separate tools chained together, not one all-in-one app: a language model for the script (Claude Opus 5.5 through OpenRouter), a multi-speaker voice model (Google's Gemini TTS), a separate speech-to-text model to catch mispronunciations (ElevenLabs Scribe), ffmpeg for mixing, an image model for cover art, and a podcast host with a real API (Transistor) for publishing. Cloudflare Workers and GitHub Actions handle scheduling and orchestration.

### What's the difference between an AI podcast and a simple text-to-speech reader?

A text-to-speech tool that reads a script aloud is a demo; an AI podcast is a production pipeline that also picks its own stories, checks its own pronunciation, and recovers from its own failures without a human editing pass. The distinction matters for reliability: a single TTS call has no fallback if a voice mispronounces a name or a story goes stale, while a full pipeline can retry, splice in a verified recording, or fall back to a default and still publish on schedule.

### Can an AI podcast like this run in languages other than English?

Yes, the same architecture works in other languages because the underlying voice model supports it. Google's Gemini TTS handles more than 70 languages and regional variants with the same style-tag controls used for English, and ElevenLabs Scribe transcribes more than 90 languages for the same pronunciation-check step. Swapping languages means changing the script model's output language and the voice IDs, not rebuilding the pipeline.

### What's the ideal episode length for a daily AI-generated podcast?

Five to fifteen minutes is the sweet spot for a daily AI-generated news podcast, long enough to cover real stories but short enough to fit a morning routine. Good Morning, Israel targets 950 to 1,150 words, which runs five to seven minutes at roughly 170 words a minute. A fixed length every day, not just a fixed schedule, is part of what makes a daily show feel like a habit rather than a one-off upload.

### How much human oversight does a fully automated AI podcast actually need?

None on a normal day: the entire episode runs from trigger to publish with nobody touching a keyboard, and the system is built to fail loudly rather than silently. If a run breaks partway through, it publishes the draft state it reached instead of stalling, and a failure automatically opens a GitHub issue so a person finds out without having to notice a missing episode first. Oversight happens after the fact, through alerts, not during the run.

### What happens if there isn't enough real news to fill an episode?

If fewer real stories qualify than the show needs, the pipeline stops the run and sends a clear alert instead of padding the episode with filler to hit the schedule. Good Morning, Israel pulls its top 8 stories from a scored list of hard news published within the last 26 hours; when that list comes up short, publishing a thin episode just to stay on time is treated as a worse outcome than skipping a day.

## Ready to see how ZipLyne can help?

[Let's build something real](https://ziplyne.agency/contact)

ZipLyne Editorial

Editorial Team

Editorial desk for ZipLyne.

---

Originally published at https://ziplyne.agency/blog/how-we-built-an-ai-podcast-that-writes-voices-and-publishes
