Sonarworks SoundID VoiceAI Review
Sonarworks’ SoundID Voice AI, launched earlier this year, is one of the most exciting new vocal plugins on the market. As singer and author Deke Sharon rightly said, “The human voice was the first instrument and remains the most powerful means of musical expression.” Building on this idea, Voice AI can transform recorded vocals into entirely new performances while preserving timing and phrasing. Beyond voice replication, it can map vocals to instruments, allowing you to turn a melody or beatboxing into a fully playable instrument line or drum sequence.
Sonarworks has made the workflow quite intuitive. The user can jump right in without the need to watch a tutorial or go through a manual. The process is simple: record or import an audio clip, capture it in VoiceAI, pick a voice model or instrument, and render a new take that preserves your timing and phrasing while replacing the voice character. It has a growing library of voice and instrument models, plus a Unison mode for quick double-tracking and choir-like effects. You can run processing locally with the perpetual license, or offload it to the cloud with the token-based option. Either way, the result is an audio clip you can print, edit, and mix like any other track. Overall, the idea is quite impressive. Let’s dive into the features and use cases even further.

VoiceAI Features
SoundID VoiceAI runs as a standard plugin inside your DAW. Once processed, you can drop the new audio clip into your project, perfectly aligned with the original timing and phrasing.
It handles up to five minutes of input audio. With the perpetual version, you can process captures endlessly, while the cloud version offers ten free re-processes per hour per preset before using tokens. Pricing starts at $10 for 1000 tokens, $45 for 5000, and $160 for 20000. Since cloud processing uses about ten tokens per second, a one-minute render costs roughly $6. Compared to the $99 perpetual license, the token model is pricey, but it’s a solid entry option—especially since it saves CPU and runs smoothly even on older machines. Both versions deliver the same quality; the main difference lies in cost and whether the processing happens locally or on Sonarworks’ servers.
Before rendering, the Voice Cleanup switch helps reduce noise and room tone, while the Transpose control adjusts the output pitch for range matching or creative use.

The model browser offers two categories: Voice and Creative. Voice models cover a range of sung and spoken tones, while Creative models map vocals to instruments, preserving phrasing and dynamics. Sonarworks also provides “Rock” and “Kids” expansion packs.
Unison mode generates multiple takes of the same output with adjustable pitch, timing, and stereo width—ideal for doubles, harmonies, or choir effects.
Overall, VoiceAI is a well-rounded tool packed with useful features. It’s excellent for demos and background layers—but is it realistic enough for lead vocals? I tested it in several scenarios to find out.
Testing and Use-Cases
VoiceAI has a range of practical uses, especially for building demos or experimenting with different vocal styles. It’s great for quick vocal stacks with Unison mode and for speeding up workflows using the voice-to-instrument function. Like most AI voice tools, full realism is still a work in progress.
I started by testing it on an older trap project—heavy 808s, snappy drums, and autotuned vocals. After capturing the audio, I processed it with several voice models to compare results. The tracking accuracy was impressive: inflections were nearly identical to the original, and timing remained perfectly aligned with no latency issues. It delivered exactly what it promised.
I tested the presets “Samuel,” “Derek,” “Elton,” “Keisha,” and “Sophia.” While the core phrasing came from the input audio, each had distinct tonal character. “Samuel” sounded chesty and nasal, “Derek” smoother and slightly pitchy, “Elton” clean and open, “Keisha” warm and even, and “Sophia” softer with a lisp on sibilants. All were clear and usable, each fitting different genres or moods—I ultimately went with “Elton.”
This has a lot of uses, particularly in creating demos or trying out different vocal styles for your song. It’s great for building quick vocal stacks with unison mode, and for speeding up your workflow with the voice-to-instrument function. Of course, as with all current AI voice tools, true realism is still a work-in-progress.
I started the test with an older project I had mixed for an artist. It was a fairly straightforward trap banger with heavy 808s, snappy drums, and smooth, autotuned vocals. After capturing the audio, I processed it on a number of different voices to try out a range of styles.
I was very impressed with how well it tracked the vocal delivery of the original vocal. The inflections were almost spot-on, near perfectly emulating the way that the artists drew out certain words or emphasised syllables. The timing was accurately aligned, too – no latency issues or mishaps in the processing. It definitely delivers on what it promises.
The presets I tried were “Samuel”, “Derek”, “Elton”, “Keisha”, and “Sophia”. There wasn’t a huge difference in anything other than the main vocal tone, as the input audio determines the pronunciation and expression of the vocal. “Samuel” had a very nasal, chesty voice, “Derek” was smoother with a pitchy twinge to the vowels, “Elton” had a very clean and clear voice, “Keisha” was similarly clean and smooth, and “Sophia” was muffled with a bit of a lisp on the sibilant sounds. They were all high quality and clearly audible, and I can imagine each being used for a different genre and style of song. I stuck with “Elton” for this track.

After mixing, most minor artifacts disappeared with some compression and EQ cleanup. In dense arrangements, VoiceAI vocals blend naturally, though it’s still less suited for tracks where vocals sit fully exposed.
Next, I tested Unison mode, layering the original vocal with two and eight voices. With pitch variation at 10–20, timing at 20–40, and width around 80, the results were a convincing doubling and choral spread. With a bit of processing and panning, the vocals sounded full and dynamic.
I also tried composing a top-line using the “Saxophone,” “Distorted Guitar,” and “Strings 2” presets. This was genuinely fun—adding pitch bends and tone shaping gave expressive, musical results.
The plugin preserves vocal texture, so singing syllables like “da da dum” introduced minor artifacts, while humming or neutral tones yielded smoother, more realistic outcomes. Matching the vocal delivery to the intended instrument helped achieve the best results, and pitch correction before rendering improved accuracy.
Unison mode can be CPU-heavy, especially at eight voices, so it’s best processed in a separate project. Rendering takes 30 seconds to two minutes depending on complexity—slower than real time, but reasonable for what it delivers. Though tweaking settings can take a few tries, the creative flexibility and time saved over recording multiple takes more than make up for it.
Beyond the basics, VoiceAI shines in creative sound design. You can warp vocals into experimental textures—perfect for background layers, ad-libs, pre-drop vox, or ambient accents. It also lets you create instruments that mirror your lead vocal melody. Using models like “Strings 1,” “Saxophone,” or “Talking Bass” can add expressive call-and-response layers or extra harmonic weight. The “Talkbox” preset, in particular, adds depth and presence to any vocal take.
You can even sketch instrumental solos by singing them. The plugin translates phrasing and dynamics so accurately that it’s a great way to mock up parts for session players—or to inspire your own solo compositions. With the right processing, these takes can sound strikingly alive.

The processing retains much of the vocal texture, so singing with clear syllables like “da da dum dum” introduced some pronunciation artifacts. Humming or using neutral tones produced smoother, more realistic results. Approximating the target sound (like a sax or guitar) in the recording made the output far more convincing. Since the plugin stays true to the input, pitch-correcting your take beforehand helps avoid off-key results.
Unison mode, especially with all eight voices, was quite CPU-heavy and best handled in a separate project. Processing times ranged from 30 seconds to about two minutes, depending on complexity—slower than instant but reasonable for what it achieves. While fine-tuning settings can feel a bit tedious, the creative flexibility and time saved over recording multiple takes make it a worthwhile trade-off.

Beyond the basic workflow, VoiceAI has some really interesting sound design applications. It can be used to pitch vocals in unnatural and strange-sounding ways – perfect for a background layer or an ad-lib style double to sprinkle across your track. This can be used for creating repetitive vox or pre-drop vocal hits in EDM tracks, or adding an otherworldly layer to ambient, spoken word songs.
Another interesting use is to create an instrument track that follows the main vocal. Using a model like “Strings 1” (for a lower-pitched viola sound), “Saxophone” or “Talking Bass” can give you an incredible supporting element to reinforce and emphasise the lead vocal. You could use this to create easy A-B sections where the instrument responds to the vocals with the exact same melody and phrasing. The “talkbox” is also a great background layer for any vocal, giving it a stronger fundamental frequency and a depth that is hard to achieve otherwise.
Instrumental solos are also notoriously difficult to translate from your head to your instrument. With VoiceAI, you can quickly create a reference for how it should sound, which should make it a lot easier to figure out in terms of timing and phrasing. This can be very helpful for working with session musicians, or even for your solo compositions – sometimes with the right processing, you can create a solo section that truly sounds alive.
Our Thoughts
Pros
- Streamlined workflow – capture, process, and print all within your DAW.
- Highly accurate timing and phrasing translation.
- Versatile library with both natural and creative models.
- Unison mode creates convincing doubles, harmonies, and choir effects.
- Voice-to-instrument translation helps with quick melody sketching and sound design.
- Cloud mode reduces CPU strain.
- Perpetual mode offers unlimited renders.
Cons
- Some AI voices sound artificial when soloed.
- Heavy CPU usage in Unison mode when processing. During Playback, CPU usage is normal.
- No custom model training or voice upload.
Is VoiceAI Worth It?
This plugin occupies a rare space in modern production. It is streamlined yet versatile, packed with vocal and instrument presets that make sketching ideas or demos remarkably fast. Beyond standard vocal processing, it opens creative pathways that show how far AI music tools have come. While it will not replace a lead singer anytime soon, it certainly empowers producers and artists to craft richer and more dynamic vocal sequences.
For hybrid producer-artists, collaborators working with vocalists, or musicians experimenting with session sounds, this is an excellent workflow enhancer. It is also simply fun to sing into your mic and transform your voice into a trumpet, guitar, or bass in seconds. Even if you are just curious, the cloud version is worth a try, and you might end up wanting the full version in your toolkit.
Price: €99.00
More info on Sonarworks.com
Also Read: