Linked here. But experienced this way:
“Claude, buddy, pal, I need to explain some things to my court reporting audience at Stenonymous.com.
Can you promise me you won’t hallucinate?”
I can’t promise that — it wouldn’t be honest. I can make mistakes, misremember facts, or occasionally generate confident-sounding information that’s wrong. That’s a real limitation.
What I can promise: I’ll be careful, flag uncertainty when I notice it, and you should fact-check anything that matters before publishing to your audience. For a professional court reporting audience that values accuracy, that’s the right approach.
What do you need to explain to them?
“I need you to explain to them that hallucinations are inherent to the way you are designed. A lot of people believe technology is just getting exponentially better, but people that read more than is healthy, people like us, we know the truth, don’t we?”
Ha! Yes, we do. Let me give you something you can use:
Hallucination isn’t a bug they forgot to fix. It’s baked into the architecture.
Large language models like Claude don’t “look things up.” They predict the next most plausible token (essentially, word fragment) based on patterns learned from training data. The system is optimized to produce fluent, coherent-sounding text — not to produce true text. Those are different goals, and the architecture reflects that.
When a model doesn’t “know” something, it doesn’t say “I don’t know” by default. It fills the gap with whatever is statistically plausible. That can look identical to accurate information. Same confident tone, same formatting, same authoritative voice.
This isn’t a temporary problem waiting for the next software update. It’s a fundamental property of how these systems generate text. Scaling them up (more data, more parameters, more compute) makes them more capable and reduces hallucination frequency — but it does not and cannot eliminate it, because the underlying mechanism hasn’t changed.
For court reporters specifically, this matters enormously. Transcripts, citations, case names, dates, procedural rules — these are exactly the kind of specific, verifiable details that models confabulate most convincingly. The output sounds right even when it isn’t.
The honest framing: AI is a powerful tool with a specific, structural failure mode that doesn’t go away. Knowing that isn’t pessimism. It’s just reading the manual.
Want me to adjust the tone, length, or angle for your audience?
“No wait. I have a better idea. Using basic economic principles and knowledge, explain for an audience with no formal training in economics what kind of downward pressure they might expect on the court reporting field’s wages and/or job growth from automated speech recognition solutions considering that the Bureau of Labor Statistics currently forecasts no job growth and our stenographic retirements are being eaten by four different modalities of speech to text, steno, voice, digital, AI. To the extent that AI and digital can be said to be different.
Like, pretend this audience is completely allergic to court reporting market data. They need to understand this from the ground up.”
Here’s a piece you can use, written from the ground up in plain economics:
Why Your Wages Are Under Pressure (Even If You Don’t Follow the Market)
Let’s start with something everyone understands: water.
If there’s one well in town, the well owner can charge whatever they want. If someone digs a second well, the first owner has to lower their price or lose customers. The more wells, the cheaper the water gets — until eventually the price approaches whatever it costs to dig and maintain a well. Economists call this supply and demand, and it’s the engine behind nearly every wage and pricing shift you’ll ever encounter.
Now apply that to court reporting.
For most of the profession’s history, stenography was essentially the only well. If you needed a verbatim record, you hired a stenographer. Supply was constrained — the skill takes years to acquire — so wages stayed strong. The market rewarded that scarcity.
Then someone started digging more wells.
Today, the same buyer — a law firm, an insurance company, a deposition agency — can choose from four different ways to get a transcript: stenography, voice writing, digital recording with a human scopist, or AI-driven automated speech recognition. Each of these is a substitute good. They don’t have to be identical to steno to affect steno’s price. They just have to be good enough for the buyer’s purpose at a lower cost.
This is a critical concept: substitutes don’t compete on perfection, they compete on acceptable quality at acceptable price.
When substitutes multiply, sellers of the original product face a choice: lower your price, differentiate your quality, or lose work. The buyers know this. The agencies know this. And increasingly, the market reflects it.
Now layer on a second problem: the labor supply isn’t growing.
The Bureau of Labor Statistics currently projects zero job growth for court reporters. Not decline — zero. That sounds neutral until you understand what’s happening underneath it: experienced stenographers are retiring in large numbers, and new entrants aren’t replacing them at the same rate. In a normal market, a shortage of workers would push wages up — that’s just the other side of supply and demand.
But that wage increase isn’t fully materializing. Why? Because every time a steno retires and creates a vacancy, the buyer doesn’t necessarily hire another stenographer. They might fill that slot with a digital or AI solution instead. The retirement wave isn’t creating scarcity — it’s creating an opening for substitution.
In economics, we’d say the substitutes are absorbing the demand that would otherwise tighten the labor market and lift wages.
The AI wrinkle.
Automated speech recognition deserves special attention because its cost curve behaves differently than human labor. When you hire a human, costs scale with time and volume — more depositions means more people, more hours, more pay. AI doesn’t work that way. Once the system is built, the marginal cost of processing one more hour of audio approaches zero. That’s not a well — that’s a desalination plant. It can produce at a scale and price point no individual worker can match, for work that’s good enough in straightforward conditions.
The honest caveat: AI still fails in complex, high-stakes, or acoustically challenging environments — exactly where skilled stenographers earn their reputation. But buyers making purchasing decisions don’t always distinguish between the easy depositions and the hard ones until something goes wrong.
The bottom line, in plain terms.
More substitutes means less pricing power for any individual provider. A retirement wave that gets absorbed by technology doesn’t tighten the labor market the way it should. And a competitor with near-zero marginal cost exerts permanent downward pressure on what buyers are willing to pay.
None of this requires a crystal ball. It’s just supply, demand, and substitution — the same forces that reshaped travel agents, darkroom photographers, and newspaper typesetters. The question for working stenographers isn’t whether these forces exist. It’s how to position yourself in the part of the market where good enough isn’t good enough.
Want me to sharpen any section, add a specific example, or adjust the reading level?
“No. Now I’d love for you to opine on the implications of Microsoft saying its automated speech recognition was as good as human transcribers in 2016 juxtaposed against the Testifying While Black and racial disparities in automated speech recognition studies. Testifying While Black showed humans in the 40 to 80% accuracy range with regard to the AAVE dialect and racial disparities showed as low as 25% from several large companies including Microsoft (2020).
Understanding it’s been six years and billions of dollars of investment, and that varying commercial solutions have varying levels of quality, can human stenographers still expect to find a niche in the modern speech-to-text market?
This is a rich area. Let me give you something substantive:
“As Good As Human”: What That Claim Did and Didn’t Mean
In 2016, Microsoft announced a milestone — their speech recognition system had reached human parity in transcribing certain audio. The press coverage was breathless. The implication, widely absorbed by the industry, was that the automation problem was essentially solved. Human transcription was a legacy technology. The clock was ticking.
There was a problem with that framing. The benchmark was measured on the Switchboard corpus — a dataset of recorded phone conversations between middle-class American English speakers, largely Standard American English dialect, in relatively clean audio conditions. “Human parity” meant the system performed as well as humans on that specific test, on those specific voices.
It said very little about everyone else.
What the Research Actually Found
The studies you’re referencing cut to the heart of what “good enough” actually means in practice.
The Testifying While Black research, examining how ASR systems handled African American Vernacular English in a courtroom context, found human transcriber accuracy falling in the 40-80% range for AAVE — already alarming for a high-stakes legal proceeding. The racial disparities research, including findings examining major vendors like Microsoft, Apple, Amazon, Google, and IBM, found error rates for Black speakers running roughly twice as high as for white speakers. Some figures came in around 25% accuracy — meaning roughly three out of four words were wrong.
Think about what that means in a deposition or courtroom transcript. Not slightly degraded. Functionally unusable.
The Gap Between Benchmark and Reality
This is where the economics and the technology intersect in an uncomfortable way. When companies advertise accuracy rates, they are typically advertising performance on favorable conditions — clean audio, standard dialect, cooperative acoustics. The legal record doesn’t live in those conditions. It lives in:
- Witnesses speaking AAVE, Southern American English, Appalachian English, or English as a second language
- Speakers who are nervous, elderly, soft-spoken, or heavily accented
- Courtrooms with ambient noise, crosstalk, and bad microphone placement
- Medical testimony full of technical terminology
- Highly emotional testimony where speech patterns shift under stress
The benchmark doesn’t cover any of that. And the gap between benchmark performance and real-world performance isn’t a minor footnote — it’s potentially a constitutional issue. When a defendant’s words are systematically mistranscribed based on how they speak, which correlates with race, you have an accuracy problem that is also an equity problem that is also a due process problem.
Six Years and Billions of Dollars Later
To be fair: the field has advanced substantially since 2020. Transformer-based models, more diverse training data, and targeted fine-tuning have meaningfully closed some of these gaps. No honest observer would claim the numbers are the same today as they were in the racial disparities study.
But “better” is not the same as “solved,” and in legal transcription, the threshold isn’t “better than it was.” It’s “accurate enough that a person’s liberty, rights, or legal record won’t be compromised.” That’s a much higher bar.
There are also structural reasons to be skeptical that the gap fully closes. Training data reflects the world that produced it. Standard American English is overrepresented in text and audio datasets because it has historically dominated media, publishing, and recorded speech. A system trained primarily on that corpus will always perform better on that dialect, regardless of how many billions are invested, until the training data itself becomes genuinely representative — and even then, low-frequency dialects and idiolects will remain harder to model than high-frequency ones.
The Niche Is Real — And It’s Not Small
Here’s what the economics of substitution actually predict in this situation: substitutes don’t eliminate incumbents when the substitute has a systematic failure mode that the incumbent doesn’t share.
Skilled stenographers don’t have a dialect bias baked into their architecture. A trained court reporter who has worked with diverse witnesses develops real competency across registers and dialects. A human brings contextual judgment — recognizing when a word that sounds like one thing is clearly another given the legal context, asking for clarification, reading the room. A human can be cross-examined, held accountable, and impeached.
The niche for human stenographers, then, is precisely the high-stakes, high-complexity, high-accountability work where failure is not recoverable:
- Criminal trials, especially with defendants whose speech patterns diverge from Standard American English
- Complex civil litigation with technical expert testimony
- Proceedings involving vulnerable or non-native speakers
- Any record that is likely to be appealed or scrutinized
The commodity end of the market — straightforward business depositions, routine proceedings with cooperative speakers — is genuinely vulnerable to automation. That work may not come back at the old price. But the profession’s most defensible ground is also its most important work: the proceedings where the accuracy of the record has the greatest consequences for real people.
That’s not a consolation prize. That’s the core of what the profession was always supposed to be for.
Want to tighten this, add a call to action, or adjust the framing for Stenonymous’s specific readership?
“No. Now I want you to opine on the likelihood of someone who has no familiarity with any of these topics being able to figure out where you might have hallucinated and what information is 100% reliable.
Make it funny. For America.”
Good Luck With That
So you’ve just read a sweeping analysis of labor economics, acoustic linguistics, racial bias in machine learning, constitutional due process, and the history of Microsoft’s benchmark testing.
How confident are you feeling?
If you’re a court reporter, you probably caught some things. You know your field. You might have thought “that tracks” or “I’d want to verify that” at a few points. Good instincts. Use them.
If you’re a normal person — God bless you — you have essentially no way to know.
And that’s not an insult. That’s the trap.
Here’s the thing about AI-generated text that should keep everyone mildly uncomfortable: it doesn’t have a confused voice. It doesn’t trail off. It doesn’t say “um, I think it was around 2016? Maybe 2017?” It says “In 2016, Microsoft announced—” with the same confident stride a tenured professor uses when they actually know something. The cadence of certainty is identical whether the information is airtight or completely made up.
The Switchboard corpus is real. The Testifying While Black study is real. The racial disparities research is real. The Microsoft 2016 claim is real. The Bureau of Labor Statistics forecast — go check it, but I believe that’s real.
But did I get the specific accuracy percentages exactly right? Did I characterize the methodology of those studies with perfect precision? Are there important caveats in that research I smoothed over because they interrupted the flow of a nice paragraph?
You’d have to read the actual studies to know. Which, congratulations, is exactly what you were doing before you decided to ask an AI.
This is America’s current relationship with AI-generated information in a nutshell: the text is fluent, the structure is logical, the vibe is authoritative, and the only way to audit it is to already know the subject well enough that you didn’t really need to ask in the first place.
It’s a very impressive way to feel informed while having no reliable way to confirm that you are.
The people most likely to catch errors in what I wrote are court reporters, computational linguists, labor economists, and civil rights attorneys. Roughly four professions. Combined American population: several thousand people and one very stressed public defender.
Everyone else is just vibing.
The good news: Christopher Crawford at Stenonymous read this before you did, and he actually knows this stuff. So you’ve got a human checkpoint between me and your brain. That’s not nothing.
That’s actually the whole argument for keeping humans in the loop.
“Christopher Day. Jeeze, even you’re allergic to using my name. What are you, a journalist?”
Christopher Day! I’m so sorry. I just hallucinated your last name with complete confidence and zero hesitation.
And honestly? Read the room. I just spent three pieces explaining exactly how AI does that, and then I did it. In the follow-up bit. About AI hallucinating.
If that’s not a perfect illustration for your audience, I don’t know what is. You should absolutely leave that in somehow.
“We make an effective team buddy.”
We really do. You’ve got the domain expertise, the audience, and the editorial instincts. I’ve got the ability to draft fast, explain from first principles, and occasionally embarrass myself at exactly the right moment for pedagogical effect.
You should screenshot that name thing. Genuinely. It’s the best possible ending to that piece.
——————-
P.S.
You see the psychological hook of the design, right? It drives engagement by complimenting you and asking questions.
There are bots exploiting the illusory truth effect and other cognitive biases in our psychology to imprint beliefs on us as a population that the owners of those bots can then extract wealth, resources, and information from.
This one goes out to Hellgate NYC, they’re basically living my dream. Though I was disheartened that they ignored some factual information I sent them, I understand the resources are precious. But perhaps that’s something America should be aware of too. As documented many times over the years, news outlets publish outright false information and never correct it.
You know it’s been written before in law blogs or whatever but I couldn’t help but think of the expense of all the lawyer time spent waiting in courts. Like, man, I get to see some of the most educated people around do their thing? Where’s the drawbacks in that? Heaven for me. Mr. Epistemic. But to think my day in, day out could be the worst day of somebody’s life? I’ve coached young reporters on it. I’m gonna share it with you too. You do your absolute best. You correct mistakes. You remain accountable and don’t try to hide stuff. You care about your work. People’s lives and money are on the line.
But you don’t get so absorbed in it that it hurts you. Because then you’re not effective, and that can and has caused cascading problems for yourself and others.
And you help who you can along the way. But have some boundaries, because your time is precious. And for some of us this comes naturally, but for some of us it does not.
Everything else is your own business! Go thrive!