Resemble AI previously made their name as one of the top players in quality voice cloning, then in late 2025 the firm took a surprising $13M turn with funding specifically for deepfake detection and audio authentication. In 2026 Resemble is the one leading voice cloning company where generating a synthesised voice and detecting an AI-generated voice are given equal status and weight. Which makes their market positioning vital for enterprise users that need both to create a synthetic voice with an organisation and guard against its nefarious use.
Chatterbox, an open-source model that debuted as part of Resemble’s pivot, can clone a voice from five seconds of audio and covers 23+ languages – one of the lowest minimum voice input requirements currently offered anywhere.
That minimum voice data input for cloning will surely have grabbed the attention of voice vendors. Along with its pivot to security, Resemble also shifted its per-minute pricing from its past plans to a pay-as-you-go system where users pay $0.03 per minute; the pricing structure is more flexible but also more costly on the downside, as it removed its previous monthly minimums. Resemble’s flagship detection product, dubbed Resemble Detect, offers 99% accuracy with synthesised voices originating from any source. Resemble’s other authentication product, PerTh watermarking technology, inscribes stealthy verification codes directly into audio upon creation.
It is worth noting a slight criticism in that by pivoting to a more secure, enterprise-oriented experience with an eye towards security, Resemble has made the platform less accessible to single-use, consumer-focused creators.
If you're looking to clone a few voices here and there for social media posts or to supplement an ebook and prefer individual workflow, you might find alternative options like ElevenLabs, currently a more streamlined approach, though this is an intentional feature shift.
- Category: Voice Cloning
- Pricing: Paid
- Rating: 4.3 / 5 (0 reviews)
- Platforms: Web
Key features
- Chatterbox open-source model — Clones voice from approximately five seconds of audio across 23 plus languages with MIT licence for self-hosting
- Five second minimum audio — One of the shortest source audio requirements for functional voice cloning in the category
- Resemble Detect — Identifies AI-generated audio with 99 percent accuracy across content from any TTS or cloning system not only Resemble-generated audio
- PerTh watermarking — Embeds inaudible authentication markers at voice generation time for future provenance verification
- Flex pay-as-you-go pricing — $0.03 per minute of generated audio with no monthly minimum commitment
- 23 plus language multilingual cloning — Generate cloned voice speech across languages from the same five-second source sample
- Real-time voice conversion — Convert live audio to a different voice in real time for streaming and call centre applications
- Enterprise API — Full programmatic access for integrating cloned voices into custom applications and pipelines
Pros & Cons
Pros
- Five second minimum audio requirement for Chatterbox is the most accessible cloning threshold in the category for quick project turnaround
- Deepfake detection and watermarking built into the same platform as voice creation is unique and increasingly valuable as audio authentication becomes a compliance concern
- Open-source Chatterbox model under MIT licence allows self-hosted deployment for teams with data sovereignty requirements
- Pay-as-you-go at $0.03 per minute avoids minimum monthly commitment for teams with variable generation volume
- $13 million enterprise security funding round reflects serious institutional belief in the combined creation-plus-detection positioning
Cons
- Enterprise and developer pivot has made consumer-facing workflows less polished than ElevenLabs for individual creators who want a clean studio experience
- Pay-as-you-go pricing is harder to budget predictably than flat monthly plans for teams with consistent high-volume generation
- Retired Creator and Professional subscription tiers mean historical plan comparisons in reviews may not reflect current pricing structure
- Real-time voice conversion quality depends on network conditions which affects reliability for production live streaming use cases
Visit Resemble AI Voice Cloning