CloudTextAndSpeech is a self-hosted voice platform that runs in both directions. It turns text into lifelike speech — audiobooks, podcasts, voice content, applications that talk — and it turns recordings into text, for transcripts of any audio, voice content or customer service call. Both halves sit behind one set of accounts, one set of plans and one admin panel, which is the difference between this and running two products side by side.
Synthesis and recognition come from four vendors on your own keys: Amazon Web Services (both directions), Google Cloud Platform (both directions), Microsoft Azure (text to speech) and IBM Cloud (text to speech). Register with one or with all four. Fully connected, that is over 900 voices across more than 144 languages and dialects for speech synthesis, and over 170 languages and dialects for transcription. Only the vendors you activate appear in the app, so starting on one and adding the rest later is a supported path rather than a workaround.
Text to speech
- Standard and Neural voices — neural (NTTS) engines for a clear step up in quality, including Google's WaveNet voices, with separate effect sets for standard and neural.
- Speaking styles — a Newscaster style tuned for news narration and a Conversational style for two-way and telephony use, where the vendor supports them.
- SSML as controls — rate, pitch and loudness, emphasis, correct pronunciation of digits, dates, abbreviations and awkward words, word and phrase replacement, and mute or beep-out over any part of a sentence. The tags available for the selected voice are listed in the app, since support varies by vendor.
- Sound Studio — mix up to 20 voices in a single synthesis task, drawn from anywhere in the 900-voice catalogue, so dialogue and multi-narrator pieces come out as one file.
- Up to 60,000 characters per task, with large text synthesisable straight into your own Amazon S3 bucket.
- Output formats — MP3 (AWS, Azure, GCP, IBM), OGG (AWS, Azure, GCP, IBM), WAV (GCP, IBM) and WEBM (Azure).
- Near real-time synthesis, with streaming audio optimisation.
Speech to text
- 170+ languages and dialects for transcription, across AWS and GCP.
- Live Transcribe — real-time transcription in 12 languages, via AWS.
- Speaker identification up to 5 people — on AWS and GCP, so interviews and calls read as a conversation.
- Instant results for short audio — GCP returns short files immediately instead of routing them through the long-file path.
- Editable live results — fix a mishearing in place rather than exporting and correcting elsewhere.
- Input formats — MP3, OGG, WEBM and MP4 (AWS), WAV and FLAC (AWS and GCP).
- File ceilings — up to 4 hours and 2 GB per file on AWS (2-channel), or up to 8 hours with no size limit on GCP (1-channel).
Storage and sharing
- Three storage targets — the local server, Amazon S3 or Wasabi, so generated audio need not accumulate on the app disk.
- Share or download — results are shareable to social media or downloadable directly.
The business layer
- Plans — monthly subscriptions and prepaid packs, both covering the whole platform rather than one half of it.
- 8 payment gateways — PayPal, Stripe, Razorpay, Paystack and Mollie for subscriptions and prepaid; Braintree and Coinbase for prepaid; and offline bank transfer for both. All of them, and every SaaS feature, are covered by the Regular License.
- Crypto — Coinbase takes Bitcoin, Bitcoin Cash, Ethereum, USD Coin, Litecoin, Dogecoin and Dai on prepaid plans.
- Coupons and promo codes — discount codes for prepaid plans.
- Affiliate and referral system — referrals, earnings and payouts, included rather than bolted on.
- Finance dashboard — monthly and yearly income against estimated spend on the cloud services, which is the pair of figures that tells you whether a plan is priced above cost.
- One-click auto update — move to a new release without a manual file drop.
Requirements worth knowing
Built on PHP 8.1 and Laravel 9, with comprehensive documentation and six months of included support. At least one cloud vendor account is required; any combination works, and reaching the full catalogue means registering with all four. Synthesis and transcription are billed by the vendors at their own published rates, which is what the estimated-spend dashboard exists to track. The interface is fully responsive, but it is a web application rather than a native mobile app.
Ideal for
- Operators who want to sell narration and transcription as one product rather than two subscriptions.
- Podcast and video teams generating voiceover and then transcribing the finished cut for show notes and captions.
- E-learning producers who need course audio in many languages and transcripts of the same material.
- Publishers narrating articles at volume and transcribing interviews for the same newsroom.
- Support and sales operations turning call recordings into text and outbound scripts into audio.
- Founders launching a paid voice platform with plans, gateways and affiliates already built.