A blog prob­a­bly of inter­est only to nerds by John F Mor­ton.

Ink of the day: MIXING… 

SuperGeekery: A blog probably of interest only to nerds by John F Morton.

Alias TTS: the audio was never the hard part

▷ Audio edition

The narration of this post was created with Bespoken plugin for Craft CMS.

The sto­ry of a long-stand­ing love of audio, an open mod­el that gave me the same jolt Eleven­Labs did the first time I heard it, and the unglam­orous qual­i­ty-con­trol lay­er that turned it into some­thing I’d actu­al­ly put my name on.

First, what is Alias TTS?

Alias TTS is a self-host­ed text-to-speech ser­vice for cre­at­ing nat­ur­al-sound­ing audio in a cloned voice. It wraps open voice mod­els in the less glam­orous machin­ery a real appli­ca­tion needs: long-form gen­er­a­tion, sen­tence-aware chunk­ing, audio cleanup, stor­age, and auto­mat­ic checks for flawed takes. It also speaks the same API lan­guages as Eleven­Labs and Ope­nAI, so an exist­ing appli­ca­tion can use it by chang­ing lit­tle more than its base URL and API key.

That’s what it is tech­ni­cal­ly. The rea­son I built it starts some­where more per­son­al.

It starts with a love of audio

I’ve been an audio per­son for a long time, and it pulls at me from two direc­tions.

One is the craft. I love mak­ing audio — the edit­ing, the pac­ing, the small choic­es that decide how a line lands. But let me be clear about some­thing up front, because it shapes every­thing else: I don’t think lis­ten­ing is bet­ter than read­ing. Read­ing is a fan­tas­tic expe­ri­ence in its own right, and I’d nev­er want to trade it away. To me the two aren’t rivals so much as com­pan­ions — par­al­lel routes to the same words, and you take whichev­er one suits the moment you’re in.

The oth­er pull is per­son­al. I get migraines. I’m also some­one who loves to read. Those two don’t get along: when a migraine set­tles in, a screen full of text is the last thing I want, but I still want the words. So I lis­ten instead — an arti­cle, a blog post, a whole book. Put it in my ears and the migraine stops stand­ing between me and the thing I want­ed to read.

Once you’ve need­ed audio that way, a lis­ten to this post” but­ton stops look­ing like a nov­el­ty. The read­ers it mat­ters most to are the ones many peo­ple would name first: those who are blind or have low vision, for whom a spo­ken ver­sion isn’t a con­ve­nience but the way in. And it reach­es fur­ther still — tired eyes, a migraine, a long com­mute, a screen some­one can’t look at right now. Along­side the writ­ing, audio lets a wider audi­ence get to the words.

Put those two pulls togeth­er — the love of mak­ing the audio and the need to have it — and you get the rea­son I built the Bespo­ken plu­g­in for Craft CMS: a way to give any arti­cle a spo­ken ver­sion, so a read­er who’d rather lis­ten — or who can’t com­fort­ably read right now — can press play and let the piece read itself to them. I released it pub­licly in Sep­tem­ber 2024, after a good stretch of build­ing it in pri­vate first, and it’s been qui­et­ly doing that job ever since.

Under the hood, Bespo­ken called out to Eleven­Labs, and land­ing there took some shop­ping around. I’d audi­tioned a stack of text-to-speech ser­vices, and at the time it gen­uine­ly wasn’t close — Eleven­Labs was the best by a wide mar­gin. That alone would’ve sold me. But the real light-bulb moment came when I tried the voice cloning: I fed it a sam­ple of my own voice and had it read a para­graph of my writ­ing back to me — in my voice, say­ing words I’d writ­ten but nev­er spo­ken aloud. I played it three times in a row. Not to check it. Just because it was a strange and won­der­ful thing to hear.

That’s the feel­ing the whole project is chas­ing.

Then I met Chatterbox

A while lat­er I came across Chat­ter­box, the open text-to-speech mod­el from Resem­ble AI. I ran a clip through it on Repli­cate, cloned my own voice from about twen­ty sec­onds of me talk­ing, and had it read a para­graph back.

There it was again — my own voice, my own words, the same jolt I’d had with Eleven­Labs. Except this time the mod­el was open, I could run it myself, and there was no month­ly ceil­ing stand­ing behind it. I sat there and lis­tened to the clip a few more times than I’ll admit to.

When it land­ed a good take, Chat­ter­box stood right next to the com­mer­cial ser­vices. That qual­i­fi­er is going to do a lot of work lat­er in this post.

The meter in the corner of my eye

Here’s a bit of fric­tion I only noticed by watch­ing my own behav­ior over time.

Eleven­Labs bills against a month­ly bal­ance of cred­its. It’s a com­plete­ly rea­son­able way to price a ser­vice, and I want to be clear that none of what fol­lows is a com­plaint about Eleven­Labs — their prod­uct is excel­lent and I still admire it.

I should also own the oth­er half of it: I’m prob­a­bly extra price-sen­si­tive here. SuperGeek­ery is a per­son­al blog. I don’t write it to make mon­ey, so every dol­lar the audio cost came straight out of my own pock­et for some­thing that was nev­er going to earn it back. Against that, a pre­mi­um rate gets hard to jus­ti­fy.

But a meter changes how you work. I noticed I’d start­ed rationing. I’d hes­i­tate to regen­er­ate a take I wasn’t quite hap­py with, because that regen­er­a­tion cost some­thing. Toward the end of a busy month I’d catch myself sit­ting on a fin­ished piece rather than spend the last of the cred­its — or bump up to the next plan for a sin­gle heavy week and then not need it again. The fric­tion was small, but it was always there, qui­et­ly nudg­ing me toward good enough” and pub­lish it lat­er.”

For a work­flow that’s inher­ent­ly iter­a­tive — gen­er­ate, lis­ten, tweak, gen­er­ate again — a finite meter is exact­ly the wrong shape.

Chat­ter­box on Repli­cate is usage-based, no sub­scrip­tion, and the num­bers aren’t close. A rough­ly 1,200-word arti­cle runs about $1.31 at Eleven­Labs’ Cre­ator-plan rate. The same text through Chat­ter­box costs me some­where around $0.18, and I only pay for what I actu­al­ly gen­er­ate. Sud­den­ly regen­er­at­ing a take cost about a pen­ny, and should I real­ly spend cred­its on this?” stopped being a ques­tion I had to ask.

The drop-in replacement that wasn’t (yet)

My first plan was the obvi­ous one: point Bespo­ken at Chat­ter­box instead of the com­mer­cial API and call it done. A drop-in swap. Same plu­g­in, cheap­er back­end, keep mov­ing.

It almost worked. And almost” turned out to be the entire project.

Chat­ter­box is non-deter­min­is­tic. Most takes were great. But every so often it would drop the last word of a sen­tence, or trail off into a strange low hum after the speech end­ed, or sim­ply stop read­ing a line part­way through. In a sin­gle short clip you just re-roll it and move on. Across a 1,200-word arti­cle stitched togeth­er from dozens of chunks, one bad chunk meant regen­er­at­ing the whole thing — and the fric­tion I’d just escaped came right back, wear­ing a dif­fer­ent hat.

That’s when the actu­al insight showed up, and it reframed every­thing:

The mod­el was nev­er the hard part. Guar­an­tee­ing the out­put was.

Access to a strong gen­er­a­tive mod­el is basi­cal­ly a com­mod­i­ty now. What didn’t exist — what I actu­al­ly need­ed — was the lay­er that turns incon­sis­tent out­put into some­thing depend­able enough to pub­lish. If I could build that, I wouldn’t have a cheap­er Eleven­Labs. I’d have the thing I’d been reach­ing for the whole time.

So in June 2026 — near­ly two years after Bespo­ken first went pub­lic — I start­ed build­ing it. Raw Chat­ter­box, straight off the open mod­el, wasn’t reli­able enough to put in front of read­ers, and it wasn’t some­thing I’d sug­gest to the hand­ful of peo­ple run­ning my plu­g­in either. If I want­ed that voice, I’d have to build the reli­a­bil­i­ty around it myself. That project became Alias TTS.

Building the thing, one failure mode at a time

What fol­lows is rough­ly the order it actu­al­ly hap­pened. The repo itself is pri­vate, but I’ve worked to keep its changel­og an hon­est record of the evo­lu­tion — it’s now well past sev­en­ty releas­es, and look­ing back, almost every ver­sion is me dis­cov­er­ing one more way audio can go sub­tly wrong.

First, stop paying for the whole article

The ear­li­est work was unglam­orous plumb­ing: split long text into short, sen­tence-aware chunks, gen­er­ate each one, and stitch them back into a sin­gle file. Chat­ter­box is a short-form mod­el, so it need­ed this any­way — but the real prize was that once each sen­tence is its own chunk, fix­ing one bad line means re-rolling one chunk, not re-ren­der­ing the whole arti­cle.

That one idea — the file is not the unit of work — qui­et­ly became the spine of the whole prod­uct.

The seams hissed

Then I actu­al­ly lis­tened to the stitched files, and near­ly every seam had a prob­lem.

Chat­ter­box, it turns out, appends a faint noise tail to a lot of its gen­er­a­tions — a lit­tle swoosh” of hiss after the speech stops. On any one clip you’d bare­ly notice. But stitch forty clips togeth­er and enough of them are swoosh­ing that the hiss lands right at the sen­tence and para­graph breaks, where the ear is most primed to hear it.

The first fix was easy: trim each clip’s tail before stitch­ing it in. Then I found longer tails — mul­ti-sec­ond low-fre­quen­cy drones. Built a detec­tor for those. Then a short, loud re-swell” blip that sat past the drone and slipped through the first detec­tor. Then a tonal swell that ramped up over a full sec­ond. Each one got its own acoustic test — zero-cross­ing rate, loud­ness rel­a­tive to the speech, how long the tail ran.

And each fix cre­at­ed its own haz­ard: cut too eager­ly and you clip a real word. A qui­et word end­ing in a soft n” or a trail­ing s” looks a lot like noise to a naïve trim­mer, and I spent more evenings than I’d like teach­ing the thing the dif­fer­ence between “…built in” and a hum. (The trick, in the end, was loud­ness: a real word tapers off, an arti­fact is loud­er than the speech around it.)

Some­where in the mid­dle of all this I built the piece that changed the shape of the prod­uct: Stu­dio. Paste in text, see exact­ly how it’ll be chun­ked, gen­er­ate the chunks, and — the impor­tant part — edit a sin­gle sen­tence and re-syn­the­size only that chunk. That’s the moment Alias stopped being a script I ran and became a tool I used.

The model just… stopped

Here’s the fail­ure mode with no tell at all: there’s noth­ing to detect.

Some­times Chat­ter­box just ends a line ear­ly. It reads The quick brown fox jumped over the” and stops. No hiss, no drone, no arti­fact — the audio is clean, it’s just miss­ing words. Every acoustic trick I’d built was use­less against it, because acousti­cal­ly noth­ing is wrong.

The only way to catch a trun­ca­tion is to lis­ten and read along. So I gave the sys­tem ears: a local Whis­per speech-recog­ni­tion side­car that tran­scribes every gen­er­at­ed take and com­pares the tran­script back to the script it was sup­posed to read. If words are miss­ing, if the take trun­cat­ed, if there’s a long stall in the mid­dle or noise where there shouldn’t be — it flags the chunk, re-rolls it auto­mat­i­cal­ly, and keeps the best valid result.

That was the turn­ing point where Alias start­ed check­ing its own work instead of trust­ing it.

Keep every take

The next idea came straight out of my back­ground in video and audio pro­duc­tion.

In a real stu­dio, you don’t record one take and ship it. You record sev­er­al, and you assem­ble the best moments. So Alias start­ed keep­ing every take — every gen­er­ate, every re-roll, every auto­mat­ic qual­i­ty re-roll — as its own immutable clip you can audi­tion, com­pare, and select. Noth­ing gets over­writ­ten. A bet­ter ear­li­er take is nev­er lost to a worse lat­er one.

Once every take is a durable, address­able asset, regen­er­ate every­thing” col­laps­es into re-roll the one line that’s wrong.” That’s the whole game.

Making final” actually final

The last piece is the one I care about most, and it comes straight out of my years in pro­duc­tion.

If you’ve ever pro­duced any­thing, you know these files. episode-final.mp3. Then episode-final-v2.mp3. Then episode-final-FINAL.mp3 — and, inevitably, episode-final-FINAL-for-real.mp3. I’ve made so many final” ver­sions of so many things that I lost count a long time ago. Final is only final until it isn’t.

So Alias lets me seal one. A fin­ished project can be sealed: Alias records the SHA – 256 of the exact approved bytes, snap­shots them, and lets me down­load a self-con­tained receipt — the final audio, a human-read­able pro­duc­tion record of what each chunk actu­al­ly said, a machine-read­able man­i­fest, and every saved take, includ­ing the reject­ed re-rolls. From then on, final” isn’t a hope­ful file­name — it’s a fin­ger­print. Any­one can lat­er drop the file onto a ver­i­fy page and con­firm it’s the untouched approved cut, with the fin­ger­print checked right in their brows­er and the audio nev­er leav­ing their device. The moment I edit any­thing, the seal clears itself, so an approved” stamp can nev­er out­live the audio it described.

Under­neath all of it sits Gen­blaze, Backblaze’s open-source Python SDK for build­ing AI media pipelines with prove­nance baked in. It dri­ves the gen­er­ate → score → re-roll → stitch loop behind one provider-agnos­tic inter­face — I wrote the adapters that plug Alias’s own engines into it — and writes a ver­i­fi­able prove­nance man­i­fest for every run. Back­blaze B2 then durably stores every take and every man­i­fest in a pri­vate buck­et. The reject­ed takes mat­ter as much as the cho­sen ones: they’re the evi­dence that a human sat there and lis­tened before sign­ing off.

Where it landed

Along the way, the drop-in replace­ment I’d giv­en up on qui­et­ly came true. Alias speaks both the Eleven­Labs and Ope­nAI TTS dialects, so an exist­ing inte­gra­tion switch­es over by chang­ing a base URL — and my Bespo­ken plu­g­in now runs on Alias in pro­duc­tion. It became the drop-in replace­ment after all, just with an entire stu­dio stand­ing behind it.

The rest filled in around that: voice cloning you can record and clean up right in the brows­er, a sec­ond faster engine with expres­sive tags like [laugh] and [sigh], a ful­ly local mode for offline devel­op­ment, pre­paid cred­it account­ing, and some­thing north of 700 auto­mat­ed tests keep­ing me hon­est.

The name took a while to land, and I sweat­ed it more than you might expect — I start­ed out in adver­tis­ing, where nam­ing a thing is part of work­ing out what it actu­al­ly is. It went through a few names as I built it (for a while it was Mim­ic,” until that turned out to col­lide with an exist­ing TTS engine, and it nev­er real­ly cap­tured the idea any­way). Alias was the one that final­ly clicked, because it means two things at once: a drop-in alter­nate end­point stand­ing in for a com­mer­cial one, and an assumed iden­ti­ty that is, under­neath, still you. Which is exact­ly what a cloned voice is.

The part that surprised me

I built Alias assum­ing its audi­ence was devel­op­ers. The API com­pat­i­bil­i­ty felt like the killer fea­ture — point your exist­ing inte­gra­tion at it, get a cheap­er and far more editable pipeline, done.

The reac­tion I didn’t expect came from friends — the non-tech­ni­cal ones. I’d shown Alias to a few of them the way you show any­one what you’ve been pour­ing your evenings into, not because I thought they’d have any use for it. But the Stu­dio clicked for them right away: gen­er­ate, com­pare, keep the best, assem­ble the final is just how any­one who’s edit­ed audio or video already thinks, API or no API. What caught me off guard wasn’t that they were polite about it — it’s that a cou­ple of them want­ed to try it them­selves. One, a doc­u­men­tary film­mak­er, has said he’d like to play around with it.

That reframed a ques­tion I’m still chew­ing on: I aimed Alias at devel­op­ers, but the peo­ple who actu­al­ly reached for it might be the ones who’d nev­er touch an API.

Where this leaves me

I’ve spent sev­en­ty-odd releas­es on Alias, and I most­ly use it for the thing I want­ed at the very start: to hear my own posts in my own voice on the days read­ing is a chore — and, late­ly, to actu­al­ly trust what I hear.

I can’t tell you what Alias real­ly is yet, or who it’s ulti­mate­ly for. And I won​’t pre­tend I’ve set­tled it. It’s a ques­tion I keep cir­cling back to — whether Alias stays just mine, or turns into some­thing I hand to oth­er peo­ple. The build­ing itself I’m not wor­ried about; I’ll keep tweak­ing and improv­ing it, because that part comes eas­i­ly to me. It’s the next step that’s hard: open­ing the gate and actu­al­ly invit­ing peo­ple into a play­ground I built for myself. That’s always been the hard­est part for me, and it’s exact­ly where I’m stand­ing now.