MVSEP Logo
  • Home
  • News
  • Plans
  • Demo
  • Tools
  • Create Account
  • Login
  • Theme
    Model Selector
    Language
    • English
    • Русский
    • 中文
    • اَلْعَرَبِيَّةُ
    • Polski
    • Portugues do Brasil
    • Español
    • 日本語
    • Français
    • Oʻzbekcha
    • Türkçe
    • हिन्दी
    • Tiếng Việt
    • Deutsch
    • 한국어
    • Bahasa Indonesia
    • Italiano
    • Svenska
    • suomi
    • български език
    • magyar nyelv
    • עִבְֿרִית
    • ภาษาไทย
    • hrvatski
    • Română
    Server DE2

Algorithms

MVSep Clavinet

The Clavinet is an electromechanical keyboard instrument developed and manufactured by the German company Hohner from 1964 to 1982. Originally designed for home use and classical music performance (as a modern electric replacement for the traditional clavichord), its unique sound made it an iconic instrument in the funk, disco, soul, and rock genres.

How It Works

In terms of sound production, the Clavinet most closely resembles an electric guitar with its mechanisms hidden inside a keyboard casing:

  • Beneath each key is a metal lever with a small rubber tip (a tangent). When a key is pressed, this tip forcefully strikes a real metal string and holds it down, causing it to vibrate.

  • The string vibrations are captured by electromagnetic pickups (a system similar to that of an electric guitar). They convert the acoustic vibrations into an electrical signal, which is then routed to an external amplifier.

  • A standard Clavinet features 60 keys (a 5-octave range).

Sound Character

The Clavinet's sound is highly recognizable. It is bright, sharp, very percussive, and snappy (staccato). Because of this dynamic response, the instrument is perfectly suited for playing fast, rhythmic, and syncopated parts.

Musicians frequently run the Clavinet's signal through guitar effects pedals. The use of a wah-wah pedal, auto-wah, or overdrive/fuzz has become a classic signature of funk music.

Famous example: The most famous Clavinet riff in music history can be heard in the intro to Stevie Wonder's "Superstition" (1972).

Main Models

Over its production run, Hohner released several modifications of the instrument:

Model Features
Clavinet C An early popular model featuring a distinctive white-and-red casing. This is the version Stevie Wonder played in the early 1970s.
Clavinet D6 The most legendary and widely used model (produced from 1971). It features a wood-grain finish, a control panel with tone switches (EQ), and a mechanical mute slider to simulate palm-muted guitar strings.
Clavinet E7 / Clavinet Duo Later versions optimized for touring and live performances. The Duo model combined the Clavinet and another electromechanical instrument—the Pianet—within a single casing.

🗎 Copy link Use algorithm Demo

MVSep Mellotron

The Mellotron is a unique electromechanical keyboard instrument developed in England in the early 1960s. Essentially, it is a distant ancestor of the modern digital sampler, but instead of digital memory, it uses strips of magnetic tape to reproduce sounds. The instrument became one of the most essential sonic elements of progressive rock and psychedelic music.

How It Works

Hidden inside the Mellotron is a complex, heavy, and somewhat temperamental mechanism:

  • Magnetic tape: Beneath each key is an individual strip of magnetic tape pre-recorded with a single note played by a real instrument (such as a violin, flute, cello, or a choir).

  • Action mechanism: When a key is pressed, a special pinch roller presses the tape against a magnetic playback head, and a motor pulls the tape through—playing the recorded note.

  • Time limit: The length of the tape for each key is limited. The maximum playback duration for a single note is about 8 seconds. If you hold the key down longer, the sound simply stops. To trigger the note again, the key must be released, allowing a spring mechanism to quickly rewind the tape to the beginning.

Sound Character

The Mellotron's sound is entirely unique and recognizable from the very first seconds. Due to mechanical imperfections, slight fluctuations in the motor's speed (wow and flutter), and the inherent traits of magnetic tape, the resulting sound is slightly unstable, "warbling," atmospheric, and melancholic.

It doesn't sound exactly like a real orchestra, but rather creates a thick, mystical, and lo-fi texture that remains highly valued by musicians today.

Famous example: The most famous Mellotron recording is the dreamy "flute" intro in The Beatles' "Strawberry Fields Forever" (1967).

Main Models

The instrument's history spans several major milestones:

Model Features
Mk I / Mk II Early studio versions from the mid-1960s. These were massive, heavy machines with two keyboards (one for rhythms and chords, and another for lead instruments).
M400 The most legendary and popular model, released in 1970. It featured a relatively compact white cabinet and a single 35-note keyboard. It allowed users to swap out tape frames to access different sets of sounds.
Mellotron Micro / M4000D Modern digital versions. Aesthetically styled after the originals, they reproduce high-quality digital samples from the original Mellotron tapes without the mechanical unreliability of the vintage units.

🗎 Copy link Use algorithm Demo

MVSep Guitar (guitar, other)

The MVSep Guitar model is based on the MDX23C, Mel Roformer and BSRoformer architectures. The model produces high-quality separation of music into a guitar part (including acoustic and electronic) and everything else. The model was compared with the Demucs4HT model (6 stems) on a guitar validation set. The metric used is SDR: the higher the better.

See the results in the table below.

Algorithm name Validation type
guitar (SDR) other (SDR)
Demucs4HT (6 stems) 5.22 12.19
mdx23c (2023.08, SDR: 4.78) 4.78 11.75
mdx23c (2024.06, SDR: 6.34) 6.34 13.31
MelRoformer (2024.06, SDR: 7.02) 7.02 13.99
BSRoformer (viperx, SDR: 7.16) 7.16 14.13
Ensemble (mdx23 + MelRoformer, SDR: 7.18) 7.18 14.15
Ensemble (BSRoformer+ MelRoformer, SDR: 7.51) 7.51 14.48
BS Roformer SW (SDR: 9.05) 9.05 16.02

 

🗎 Copy link Use algorithm Demo

MVSep Pedal Steel Guitar

Pedal steel guitar is a type of electric guitar mounted on a special horizontal stand (console). The musician plays it while seated, using not only their hands, but also their feet and knees.

Key features of the instrument:

  • Pedal and lever mechanism: This is the instrument’s main distinguishing feature. A system of foot pedals and knee levers mechanically changes the tension (and pitch) of specific strings during performance. This allows the player to perform complex harmonic transitions and change chords without moving the hand along the neck.

  • Playing technique: A pedal steel guitar has no traditional frets to press with the fingers. With the left hand, the guitarist smoothly glides a heavy metal bar (slide or “steel”) across the strings, creating the characteristic sliding effect (glissando). The right hand plucks the strings using a plastic thumb pick and metal finger picks on the remaining fingers.

  • Necks and strings: Professional models often feature two parallel necks with different tunings (for example, E9 for country music and C6 for jazz and swing). Each neck usually has 10 or 12 strings.

The instrument is famous for its sustained, singing, and often “crying” tone with very long sustain. Historically and culturally, the pedal steel guitar is strongly associated with country music and western swing. However, thanks to its unique expressive capabilities, today it can be heard in a wide variety of genres, from jazz and Hawaiian music to ambient, indie rock, and pop music.

🗎 Copy link Use algorithm Demo

MVSep Plucked Strings (plucked-strings, other)

The MVSep Plucked Strings is a high quality model for separating music into plucked string instruments and everything else. List of instruments: ukulele, sitar, banjo, mandolin, dobro, harp, guitar, acoustic guitar, electric guitar.

🗎 Copy link Use algorithm Demo

MVSep Bowed Strings (strings, other)

The MVSep Bowed Strings is a high quality model for separating music into bowed string instruments and everything else. List of instruments: Fiddle, Violin, Viola, Cello, Double Bass.

🗎 Copy link Use algorithm Demo

MVSep Wind (wind, other)

The MVSep Wind model produces high-quality separation of music into a wind part and everything else. The MVSep Wind model exists in 2 different variants based on following architectures: MelRoformer and SCNet Large. Wind includes 2 categories of instruments: brass and woodwind. More specific we inluded in wind: flute, saxophone, trumpet, trombone, horn, clarinet, oboe, harmonica, bagpipes, bassoon, tuba, kazoo, piccolo, flugelhorn, ocarina, shakuhachi, melodica, reeds, didgeridoo, mussette, gaida.

Quality metrics

Algorithm name Wind dataset
SDR Wind SDR Other
MelBand Roformer 6.73 16.10
SCNet Large 6.76 16.13
MelBand + SCNet Ensemble 7.22 16.59
MelBand + SCNet Ensemble (+extract from Instrumental) --- ---
BS Roformer 9.82 19.19

🗎 Copy link Use algorithm Demo

MVSep Brass (brass, other)

The MVSep Brass is a high quality model for separating music into brass wind instruments and everything else. List of instruments: trumpet, trombone, horn, tuba, flugelhorn, untagged brass.

🗎 Copy link Use algorithm Demo

MVSep Woodwind (woodwind, other)

The MVSep Woodwind is a high quality model for separating music into woodwind instruments and everything else. List of instruments: oboe, saxophone, flute, bassoon, clarinet, piccolo, english horn, untagged woodwind.

🗎 Copy link Use algorithm Demo

MVSep Bagpipes (bagpipes , other)

The bagpipe (Bagpipes) is a traditional wind musical instrument known for its characteristic piercing and continuous sound.

How it is constructed:

  • The bag (reservoir): Usually made of animal skin or modern synthetic materials. It serves to store air.

  • Blowpipe: Through this, the musician fills the bag with air using their mouth (in some variations, small bellows pumped by the elbow are used instead).

  • Melody pipe (chanter): A pipe with finger holes, on which the musician plays the main melody by moving their fingers.

  • Drone pipes (drones): One or more pipes that produce a constant, sustained background chord on a single note.

The principle of playing is that the musician inflates the bag and then presses on it with their arm, evenly pushing air into the sound pipes. Thanks to this reservoir, the music does not stop, even when the performer takes a breath.

Although the bagpipe is most often associated with Scotland (Great Highland Bagpipe) and Celtic culture, various historical variations of it exist throughout Europe, North Africa, and the Middle East.

🗎 Copy link Use algorithm Demo

MVSep Percussion (percussion, other)

The MVSep Percussion is a high quality model for separating music into percussion instruments and everything else. List of instruments: bells, tubular bell, cow bell, congas, celeste, marimba, glockenspiel, tambourine, timpani, triangle, wind chimes, bongos, clap, xylophone, mallets, metal bars, wooden bars.

🗎 Copy link Use algorithm Demo

BandIt Plus (speech, music, effects)

BandIt Plus model for separating tracks into speech, music and effects. The model can be useful for television or film clips. The model was prepared by the authors of the article "A Generalized Bandsplit Neural Network for Cinematic Audio Source Separation" in the repository on GitHub. The model was trained on the Divide and Remaster (DnR) dataset. And at the moment it has the best quality metrics among similar models.

Quality table

Algorithm name DnR dataset (test)
SDR Speech SDR Music SDR Effects
BandIt Plus 15.62 9.21 9.69
🗎 Copy link Use algorithm Demo

BandIt v2 (speech, music, effects)

Bandit v2 is a model for cinematic audio source separation in 3 stems: speech, music, effects/sfx. It was trained on DnR v3 dataset.

More information in official repository: https://github.com/kwatcharasupat/bandit-v2
Paper: https://arxiv.org/pdf/2407.07275

🗎 Copy link Use algorithm Demo

MVSep DnR v3 (speech, music, effects)

MVSep DnR v3 is a cinematic model for splitting tracks into 3 stems: music, sfx and speech. It is trained on a huge multilingual dataset DnR v3. The quality metrics on the test data turned out to be better than those of a similar multilingual model Bandit v2. The model is available in 3 variants: based on SCNet, MelBand Roformer architectures, and an ensemble of these two models. See the table below:

Algorithm name SDR Metric on DnR v3 leaderboard
music (SDR) sfx (SDR) speech (SDR)
SCNet Large  9.94 11.35 12.59
Mel Band Roformer 9.45 11.24 12.27
Ensemble (Mel + SCNet) 10.15 11.67 12.81
Bandit v2 (for reference) 9.06 10.82 12.29
🗎 Copy link Use algorithm Demo

MVSep Braam

Braam is not a traditional physical instrument, but a powerful cinematic sound effect (virtual instrument) that has become an absolute standard in modern film and trailer music.

Main features:

  • Sound: It is a massive, low-frequency, rumbling, and often aggressive sound. It resembles an apocalyptic blast of a huge ship's horn, heavy metallic scraping, or an alarm signal.

  • Origin: This sound gained massive popularity after the release of the movie "Inception" (2010) with music by Hans Zimmer, which is why it is often called the Inception Horn.

  • How it is created: As a rule, it is the result of complex sound design. The base is formed by powerful low brass instruments (trombones, tubas, French horns). Then they are layered over heavy synthesizer basses and heavily processed with effects: distortion, saturation, and deep reverberation.

Today, Braam exists in the form of ready-made samples and libraries for virtual synthesizers (VST plugins), which composers use to instantly give a track scale, tension, or an epic feel.

🗎 Copy link Use algorithm Demo

MVSep Risers

Risers stem is an audio track containing transitional sound effects designed to build tension and energy before an important musical moment, such as a drop, chorus, or climax. Risers typically feature gradually increasing pitch, volume, density, or noise intensity, creating a sense of anticipation and momentum. These sounds are widely used in electronic music, cinematic sound design, trailers, pop music, and game audio. Common examples include synthesized sweeps, reversed cymbals, noise buildups, tonal uplifters, and layered atmospheric effects.

🗎 Copy link Use algorithm Demo

Apollo Enhancers (by JusperLee, Lew, baicai1145)

The algorithm restores the quality of audio. Model was proposed in this paper and published on github.

There are 3 models available:
1) MP3 Enhancer (by JusperLee) - it restores MP3 files compressed with bitrate 32 kbps up to 128 kbps. It will not work for files with larger bitrate.
2) Universal Super Resolution (by Lew) - it restore higher frequences for any music
3) Vocals Super Resolution (by Lew) - it restore higher frequences and overall quality for any vocals

🗎 Copy link Use algorithm Demo

Reverb Removal (noreverb)

Set of different models to remove reverberation effect from music/vocals.

Author Architecture Works with SDR (no independent testing yet) Link
FoxJoy MDX-B Full track ~6.50  
aufr33 and jarredou MDX23C Full track --- Github
anvuew MelRoformer Only vocals 7.56  
anvuew BSRoformer Only vocals 8.07  
anvuew v2 MelRoformer Only vocals ---  
Sucial MelRoformer Only vocals 10.01  
anvuew BSRoformer Only vocals (Room) 13.74 HF Link
anvuew BSRoformer Only vocals (Stereo) 22.50 HF Link

Test on our internal reverberation dataset below:

Model only vocals vocals drums bass other several
Reverb removal by FoxJoy (MDX-B) 1.2938 8.2146 5.0743 7.2590 8.0154 4.2456
Reverb removal by aufr33 and jarredou (MDX23C) 0.9761 7.3888 4.0913 5.8021 7.7194 3.4537
Reverb removal by anvuew (MelRoformer) 2.3110 2.3029 2.2408 1.8141 2.9739 1.8177
Reverb removal by anvuew (BSRoformer) 2.1902 1.4094 1.4903 1.3958 2.0425 1.2422
Reverb removal by anvuew v2 (MelRoformer) 3.4083 2.3706 1.8884 1.9344 2.6079 1.7384
Reverb removal by Sucial (MelRoformer) 0.1599 0.1750 0.8917 0.9148 0.9803 0.5664
Reverb removal by Sucial v2 (MelRoformer) 0.2052 0.7266 0.9363 --- 1.5508 0.7340
DeReverb room by anvuew (BSRoformer) 2.6593 3.0581 0.0887 1.6156 3.4134 -13.7106
DeReverb stereo by anvuew (BSRoformer) 4.3740 5.3489 5.0900 4.2709 5.3950 4.3072
Reference (SDR between reverb and orig stem) -3.38 4.01 2.35 4.47 4.65 1.04

The test was prepared based on the MUSDB18-HQ test set (50 tracks). First, all stems were dereverberated using the FoxJoy model. Then, we generated 6 different test scenarios:

  • Only vocals: Reverb was applied exclusively to the vocal stem, and only this stem was included in the mixture.

  • Vocals: Reverb was applied to the vocal stem, which was then combined with all the other stems to form the mixture.

  • Drums: Reverb was applied to the drum stem, which was then combined with all the other stems to form the mixture.

  • Bass: Reverb was applied to the bass stem, which was then combined with all the other stems to form the mixture.

  • Other: Reverb was applied to the "other" stem, which was then combined with all the other stems to form the mixture.

  • Several: Reverb was applied to multiple stems within the track, which were all combined to form the mixture.

Validation dataset 80 tracks: reverb for all stems

Model sdr bleedless fullness  l1_freq
Reverb removal by FoxJoy (MDX-B) --- --- --- ---
Reverb removal by anvuew (BSRoformer) 0.89 27.33 4.85 15.03
DeReverb stereo by anvuew (BSRoformer) 4.28 27.29 13.10 23.14
Reverb removal by aufr33 and jarredou (MDX23C) 4.81 23.17 18.65 24.92
Reverb removal by anvuew (MelRoformer) 1.86 28.93 6.62 17.17
Reverb removal by anvuew v2 (MelRoformer) 1.98 30.61 6.86 17.83
Reverb removal by Sucial (MelRoformer) 0.67 40.76 3.98 14.23
Reverb removal by Sucial v2 (MelRoformer) 0.73 35.88 4.28 14.35
DeReverb room by anvuew (BSRoformer) --- --- --- ---
MVSep Team (BSRoformer) 7.87 28.06 26.12 32.33

Validation dataset 27 tracks: reverb for vocals only

Model sdr bleedless fullness l1_freq
Reverb removal by FoxJoy (MDX-B) --- --- --- ---
Reverb removal by anvuew (BSRoformer) 6.67 27.68 13.25 32.32
DeReverb stereo by anvuew (BSRoformer) 8.80 32.84 16.50 37.48
Reverb removal by aufr33 and jarredou (MDX23C) 6.18 16.04 19.32 31.93
Reverb removal by anvuew (MelRoformer) 6.63 29.67 13.14 32.02
Reverb removal by anvuew v2 (MelRoformer) 8.20 29.94 15.91 36.13
Reverb removal by Sucial (MelRoformer) 4.88 31.56 11.13 27.90
Reverb removal by Sucial v2 (MelRoformer) 4.94 27.85 11.39 28.14
DeReverb room by anvuew (BSRoformer) --- --- --- ---
MVSep Team (BSRoformer) 9.81 29.23 18.56 40.33

Reverberation (Reverb) is the physical process of gradual sound decay in an enclosed space after the sound source has stopped. If a regular echo is distinct, separate copies of a sound (like shouting in the mountains: "Hello... hello... hello"), then reverberation — is a dense, continuous humming cloud of thousands of blended reflections from walls, the floor, the ceiling, and other surfaces (like the sound of a clap in an empty cathedral or a stairwell).

In audio engineering, the reverb effect is used to place a dry (studio-recorded) sound into a virtual space and give it volume and depth.

What does reverberation consist of?

Acoustically, this process can be divided into three stages:

  1. Direct Sound: The sound wave that reaches the listener or microphone in a straight line, without any reflections. This is the loudest and clearest signal.

  2. Early Reflections: The first echoes that bounce off the nearest surfaces and reach the ears a few milliseconds after the direct sound. They are what give our brain information about the size and shape of the room we are in.

  3. Reverb Tail (Late Reflections): A multitude of chaotic, intertwining reflections that bounce off surfaces again and again. They merge into a continuous hum and gradually lose energy (decay).

Main parameters in reverb plugins

When you open a reverb plugin in a DAW (Digital Audio Workstation), you control the physical properties of this virtual room:

  • Size / Room Size: Sets the volume of the virtual space (from a tiny vocal booth to a massive stadium).

  • Decay / Reverb Time / RT60: The time (usually in seconds) it takes for the reverb tail to decay by 60 decibels, meaning it practically disappears.

  • Pre-Delay: A very important parameter that sets the pause (in milliseconds) between the direct sound and the onset of reverberation. Increasing Pre-Delay helps separate the vocal or instrument from the "tail", preserving their clarity while maintaining the sense of a large space.

  • Damping: Simulates sound absorption. In real life, soft surfaces (carpets, people, curtains) quickly absorb high frequencies, so a long reverb tail usually sounds more muffled than the direct signal.

  • Mix / Dry/Wet: The ratio between the original dry signal (Dry) and the processed signal (Wet).

Why is reverb necessary in music mixing?

  • Creating depth (staging): Reverb acts as the Z-axis (depth) in a mix. A loud and dry sound appears close to the listener (right in their face), while a quiet sound with a lot of reverb seems distant.

  • Gluing the mix: If all instruments are recorded in different deadened studios, the mix can sound disjointed. Sending them to a common reverb bus (even in small amounts) places them into a single acoustic space.

  • Artistic effect: Creating an ethereal, ambient, or epic atmosphere (for example, the Shimmer effect, where the reverb tail is also pitched up an octave).

Why is it necessary to remove the reverb effect?

Removing reverberation (or dereverberation) — is the process of cleaning an audio signal from acoustic room reflections to obtain the original dry sound. Although reverb makes a sound beautiful and spacious, in many professional scenarios, this effect turns into unwanted noise or a serious obstacle. Here are the main reasons why there is a need to "dry out" the sound:

  • Music Source Separation: When extracting vocals or individual instruments from a mixed stereo track, reverb tails create a serious problem — they "eat into" the useful signal. Effective dereverberation allows you to get a truly clean acapella or instrument stem that sounds as if it was just recorded in a studio, rather than extracted from a concert hall.

  • Automatic Speech Recognition (ASR) systems: Room echo and hum are the worst enemies of acoustic models. Reflections "smear" short consonant sounds and phonemes. In complex machine learning tasks, such as creating children's speech recognition models, where articulation is often already unclear, the presence of reverberation catastrophically reduces transcription accuracy. Therefore, dereverberation is a critical preprocessing step for audio datasets.

  • Sampling and remixing: If you take a vocal sample or a drum loop from an old recording, it already contains the space of the original mix. If you add this sample to your track and apply your own new reverb on top of it, it will result in acoustic "mud" (the effect of reverb on reverb). To integrate someone else's sound into your mix architecture, it must first be cleaned.

  • Video and film post-production (ADR & Location Sound): Actors' speech is often recorded with shotgun microphones right on the set (for example, in an echoey empty room or a stairwell). For the dialogue to sound tight, intelligible, and studio-quality, the sound engineer needs to suppress the natural reflections of the location.

  • Restoration and forensics: Recordings from surveillance cameras, hidden microphones, or dictaphones often contain so much room hum that the words become unintelligible. Suppressing the reverberation helps restore speech intelligibility.

How does it work technologically? In the past, sound engineers tried to combat the room using Noise Gates and Transient Shapers, which simply cut off the quiet tails of the sounds. This was a crude method and often distorted the useful signal itself. Today, the task of dereverberation is solved using AI and neural networks that are trained to analyze the spectrogram, distinguish direct signal patterns from reflection patterns, and mathematically subtract the latter without damaging the original.

🗎 Copy link Use algorithm Demo

AudioSR (Super Resolution)

Algorithm AudioSR: Versatile Audio Super-resolution at Scale. Algorithm restores high frequencies. It works on all types of audio (e.g., music, speech, dog, raining, ...). It was initially trained for mono audio, so it can give not so stable result on stereo.

Metric on Super Resolution Checker for Music Leaderboard (Restored): 25.3195
Authors' paper: https://arxiv.org/pdf/2309.07314
Original repository: https://github.com/haoheliu/versatile_audio_super_resolution
Original inference script prepared by @jarredou: https://github.com/jarredou/AudioSR-Colab-Fork

🗎 Copy link Use algorithm Demo

FlashSR (Super Resolution)

FlashSR - audio super resolution algorithm for restoring high frequencies. It's based on paper FlashSR: One-step Versatile Audio Super-resolution via Diffusion Distillation. 

Metric on Super Resolution Checker for Music Leaderboard (Restored): 22.1397
Original repository: https://github.com/jakeoneijk/FlashSR_Inference
Inference script by @jarredou: https://github.com/jarredou/FlashSR-Colab-Inference

🗎 Copy link Use algorithm Demo

Stable Audio Open Gen

Audio generation based on a given text prompt. The generation uses the Stable Audio Open 1.0 model. Audio is generated in Stereo format with a sample rate of 44.1 kHz and duration up to 47 seconds. The quality is quite high. It's better to make prompts in English.

Example prompts:
1) Sound effects generation: cats meow, lion roar, dog bark
2) Sample generation: 128 BPM tech house drum loop
3) Specific instrument generation: A Coltrane-style jazz solo: fast, chaotic passages (200 BPM), with piercing saxophone screams and sharp dynamic changes

🗎 Copy link Use algorithm Demo

Whisper (extract text from audio)

Whisper is a pre-trained model for automatic speech recognition (ASR) and speech translation. It has several version. On MVSep we use the largest and the most precise: "Whisper large-v3". The Whisper large-v3 model was trained on several millions hours of audio. It's multilingual model and it guesses the language automatically. To apply model to your audio you have 2 options: 
1) "Apply to original file" - it means that whisper model will be applied directly to file you submit
2) "Extract vocals first" - in this case before using whisper, BS Roformer model is applied to extract vocals first. It can remove unnecessary noise to make output of Whisper better.

Original model has some problem with transcription timings. It was fixed by @linto-ai. His transcription is used by default (Option: New timestamped). You can return to original timings by choosing option "Old by whisper".

More info on model can be found here: https://huggingface.co/openai/whisper-large-v3 and here: https://github.com/openai/whisper

🗎 Copy link Use algorithm Demo

Parakeet (extract text from audio)

Parakeet is a family of state-of-the-art Automatic Speech Recognition (ASR) models developed by NVIDIA in collaboration with Suno.ai. These models are built on the Fast Conformer architecture, designed to deliver a balance of high transcription accuracy and exceptional inference speed. They are widely recognized for outperforming much larger models (like OpenAI's Whisper) in efficiency while maintaining competitive or superior Word Error Rates (WER). Quality metric WER: 6.03 on Huggingface Open ASR Leaderboard.

MVSep provide two versions of model (v2 and v3):
Model page v2: https://huggingface.co/nvidia/parakeet-tdt-0.6b-v2
Model page v3: https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3


Parakeet v2 (Parakeet TDT 0.6B v2)

Released as a highly efficient English-focused model, v2 established Parakeet as a leader in speed-to-accuracy ratio.

  • Language: English (en-US) only.
  • Size: 0.6 Billion parameters (600M), making it lightweight compared to the 1.1B parameters of previous versions.
  • Performance: It achieves industry-leading accuracy (approx. 6% WER on standard benchmarks) and is noted for being up to 50x faster than real-time.
  • Capabilities:
    • Supports highly accurate word-level timestamps.
    • Includes automatic punctuation and capitalization.
    • Effective at transcribing non-speech sounds like music lyrics and spoken numbers.
    • Can handle long-form audio (up to 11 hours in some configurations) using local attention mechanisms.

Parakeet v3 (Parakeet TDT 0.6B v3)

The v3 release marked the expansion of the efficient Parakeet architecture from English-only to a multilingual domain without increasing the model size.

  • Language: Multilingual (supprts 25 Euoropean languages, including English, Spanish, French, German, Russian, and others).
  • Size: Retains the compact 0.6 Billion parameter size.
  • Key Upgrade: It is trained on the massive Granary multilingual corpus (approx. 1 million hours of audio).
  • New Features:
    • Automatic Language Detection: The model can identify the spoken language from the audio signal and transcribe it without manual prompting.
    • High Throughput: Despite the added multilingual capabilities, it retains the ultra-fast inference speeds of the v2 TDT architecture.
    • Versatility: It serves as a drop-in replacement for v2 for users requiring support for European languages while maintaining low latency and compute costs.

🗎 Copy link Use algorithm Demo

VibeVoice (Voice Cloning)

VibeVoice — is a model for generating natural conversational dialogues from text with the ability to use a reference voice for cloning purposes.

Key features:

  • Two models: small and large
  • Up to 90 minutes of generated audio
  • Language support: 2 languages are supported: English (default) and Chinese
  • Voice cloning: ability to upload a reference audio recording

How to use the model

  • The text must be only in English or Chinese; quality is not guaranteed for other languages. Maximum text length is 5000 characters. Avoid special characters. 
  • Audio with the reference voice requires 5 to 15 seconds. If your track is longer, it will be automatically trimmed at the 15th second. 
  • The reference track should contain only voice and nothing else. If you have background sounds or music, use the "Extract vocals first" option.

How to generate a reference track?

We need phonetic diversity (all sounds of the language) and lively intonation. A text length of about 35–40 words when read calmly will take just ~15 seconds.

Here are three options in English for different tasks:

Option 1: Universal (Balanced & Clear)

The best choice for general use. Contains complex sound combinations to tune clarity.

"To create a perfect voice clone, the AI needs to hear a full range of phonetic sounds. I am speaking clearly, taking small pauses, and asking: can you hear every detail? This short sample captures the unique texture and tone of my voice."

Option 2: Conversational (Vlog & Social Media)

For voiceovers in videos, YouTube, or blogs. Read vividly, with a smile, changing the pitch of your voice.

"Hey! I’m recording this clip to test how well the new technology works. The secret is to relax and speak exactly like I would to a friend. Do you think the AI can really copy my style and energy in just fifteen seconds?"

Option 3: Professional (Business & Narration)

For presentations, audiobooks, or official announcements. Read confidently, slightly slower, emphasizing word endings.

"Voice synthesis technology is rapidly changing how we communicate in the digital age. It is essential to speak with confidence and precision to ensure high-quality output. This brief recording provides all the necessary data for a professional and accurate digital clone."


Tips for recording:

  1. Pronunciation: Try to articulate word endings clearly (especially t, d, s, ing). Models "love" clear articulation.

  2. Flow: Don't read like a robot. In English, melody (voice melody) is important — the voice should "float" up and down a bit, rather than sounding on a single note.

  3. Breathing: If you pause at a comma or period, don't be afraid to take an audible breath. This will add realism to the clone.

🗎 Copy link Use algorithm Demo

VibeVoice (TTS)

VibeVoice (TTS) — is a model for generating natural conversational dialogues from text, capable of creating dialogues with up to 4 speakers and durations of up to 90 minutes.

Key Features:

  • Two models: small and large
  • Up to 4 speakers in a single recording
  • Up to 90 minutes of generated audio
  • Language support: officially supports 2 languages: English (default) and Chinese, but it has been verified to work decently for other languages as well.

How to use the model

The text must be in English or Chinese; quality is not guaranteed for other languages. The maximum text length is 5000 characters. Avoid special characters. The text must be formatted specifically to indicate speakers:

Correct format:

Speaker 1: Hello! How are you today?
Speaker 2: I'm doing great, thanks for asking!
Speaker 1: That's wonderful to hear.
Speaker 3: Hey everyone, sorry I'm late!

Incorrect format:

Hello! How are you today?
I'm doing great!

Important:

  • Each line must start with Speaker N: (where N is a number from 1 to 4)
  • Speaker numbering: Speaker 1, Speaker 2, Speaker 3, Speaker 4
  • You can use from 1 to 4 speakers
  • Case does not matter: Speaker 1: = speaker 1: = SPEAKER 1

If you need a monologue, you do not need to specify a speaker.

Example scenarios:

Monologue (1 speaker):

Speaker 1: Today I want to talk about artificial intelligence.
Speaker 1: It's changing our world in incredible ways.
Speaker 1: From healthcare to entertainment, AI is everywhere.

Dialogue (2 speakers):

Speaker 1: Have you tried the new restaurant downtown?
Speaker 2: Not yet, but I've heard great things about it!
Speaker 1: We should go there this weekend.
Speaker 2: That sounds like a perfect plan!

Group conversation (3-4 speakers):

Speaker 1: Welcome to our podcast, everyone!
Speaker 2: Thanks for having us!
Speaker 3: It's great to be here.
Speaker 4: I'm excited to share our thoughts today.
Speaker 1: Let's start with introductions.
🗎 Copy link Use algorithm Demo

Qwen3-TTS (Custom Voice)

Qwen3-TTS is a powerful speech generation model offering comprehensive support for voice cloning, voice design, ultra-high-quality human-like speech generation, and natural language-based voice control. It provides developers and users with the most extensive set of speech generation features available. At MVSep, we use the largest 1.7 billion parameter model.

Original model page: https://github.com/QwenLM/Qwen3-TTS

Qwen3-TTS (Custom Voice) offers a set of 9 pre-defined speakers. Optionally, you can specify a "Voice description" to include emotions like "happy voice" or "sad voice". You can also choose the language for this model or leave it as "auto".

🗎 Copy link Use algorithm Demo

Qwen3-TTS (Voice Design)

Qwen3-TTS is a powerful speech generation model offering support for voice cloning, voice design, ultra-high-quality human-like speech generation, and natural language-based voice control. It provides developers and users with the most extensive set of speech generation features available. At MVSep, we use the largest 1.7 billion parameter model.

Original model page: https://github.com/QwenLM/Qwen3-TTS

Qwen3-TTS (Voice Design) allows you to generate speech with a custom voice that can be described in detail in the "Voice description" field. You can specify the speaker's gender and age, and add emotions, such as "happy voice" or "sad voice". You can also choose the language for this model or leave it as "auto".

🗎 Copy link Use algorithm Demo

Qwen3-TTS (Voice Cloning)

Qwen3-TTS is a powerful speech generation model offering support for voice cloning, voice design, ultra-high-quality human-like speech generation, and natural language-based voice control. It provides developers and users with the most extensive set of speech generation features available. At MVSep, we use the largest 1.7 billion parameter model.

Original model page: https://github.com/QwenLM/Qwen3-TTS

Qwen3-TTS (Voice Cloning) allows you to upload a reference audio file to generate the target text using the sample voice. To improve cloning quality, you can optionally provide the audio transcript in the "Reference text in audio" field. You can also choose the language for this model or leave it as "auto".

🗎 Copy link Use algorithm Demo

Mega 53-stem Model

Supported instruments: accordion, acoustic-guitar, back-vocal, banjo, bass, bassoon, bells, bowed_strings, brass, cello, clarinet, congas, digital-piano, dobro, double-bass, drums, electric-guitar, flute, french-horn, glockenspiel, guitar, harmonica, harp, harpsichord, hh, keys, kick, lead-vocal, mandolin, marimba, oboe, organ, percussion, piano, saxophone, sitar, snare, strings, synth, tambourine, timpani, toms, triangle, trombone, trumpet, tuba, ukulele, viola, violin, vocals, wind, wind-chimes, woodwind

Note 1: The model outputs only those instruments that were detected in the musical composition. Instruments that are not present in the track are not included in the output.

Note 2: Individual models for each instrument generally produce better results than this multimodel. Therefore, it is recommended to use this model to determine the set of stems first, and then extract them using separate instrument-specific models trained on narrower tasks.

Note 3: To reduce disk space usage, results are saved in all formats except WAV (FLAC is used instead of WAV).

Note 4: This model differs from our previously published open-source model and is an improved version of it. The comparison table below is provided:

Instrument Open-source model SDR New model SDR Delta SDR
accordion 6,2494 6,6498 +0,4004
acoustic-guitar 5,0024 5,0797 +0,0773
back-vocal 6,4179 6,5262 +0,1083
banjo 3,1593 3,7532 +0,5939
bass 11,1680 11,2886 +0,1206
bassoon 4,6595 5,1663 +0,5068
bells 1,1190 4,8040 +3,6850
bowed_strings 12,4486 12,4486 0,0000
brass 6,7042 6,8487 +0,1445
cello 5,0364 5,2257 +0,1893
clarinet 5,0505 5,8690 +0,8185
congas 9,1747 9,5946 +0,4199
digital-piano 7,9634 9,0179 +1,0545
dobro 7,6562 8,2290 +0,5728
double-bass 14,1731 15,6032 +1,4301
drums 9,2502 11,4520 +2,2018
electric-guitar 8,1543 8,1856 +0,0313
flute 6,2134 6,8557 +0,6423
french-horn 5,2635 5,5136 +0,2501
glockenspiel 3,6621 8,1523 +4,4902
guitar 2,5661 2,6164 +0,0503
harmonica 10,9265 11,8575 +0,9310
harp 6,3767 7,6523 +1,2756
harpsichord 1,6090 1,9524 +0,3434
hh 2,4681 3,2904 +0,8223
keys 9,3032 9,3106 +0,0074
kick 11,5485 11,5485 0,0000
lead-vocal 5.4663 5.4663 0,0000
mandolin 4,4256 4,7735 +0,3479
marimba 4,4821 4,8830 +0,4009
oboe 3,3616 4,5555 +1,1939
organ 10,3684 10,8244 +0,4560
percussion 2,5008 2,8897 +0,3889
piano 6,7787 6,8080 +0,0293
saxophone 8,8875 9,4985 +0,6110
sitar 4,6529 5,0434 +0,3905
snare 6,1338 6,7784 +0,6446
strings 8,8151 8,8151 0,0000
synth 2,0539 2,0539 0,0000
tambourine 3,0589 3,5544 +0,4955
timpani 4,6423 4,9779 +0,3356
toms -2,0607 -1,0708 +0,9899
triangle 5,9274 6,0197 +0,0923
trombone 2,6949 3,0751 +0,3802
trumpet 4,9658 5,8668 +0,9010
tuba 7,1957 7,5229 +0,3272
ukulele 6,7869 6,9721 +0,1852
viola 1,8581 1,8581 0,0000
violin 3,1018 3,3285 +0,2267
vocal 11,6590 11,6590 0,0000
wind 8,6317 8,6632 +0,0315
wind-chimes 4,6440 6,7529 +2,1089
woodwind 3,3123 3,3224 +0,0101
🗎 Copy link Use algorithm Demo

Upmixing by Jazzpear

This neural network model is designed for the high-quality conversion of stereo audio into spatial surround sound. Optimized for music and cinematic content (movies, TV shows, anime), the algorithm internally separates the original stereo signal into four independent stems. The platform then automatically combines these stems into a ready-to-use 5.1 multichannel file.

Stem Separation Overview

The algorithm processes the stereo file and generates four foundational tracks:

  • LR (Front): Front stereo pair

  • S (Sides): Side/surround stereo pair

  • LFE (Sub): Low-frequency effects channel

  • C (Center): Center front channel

Technology & Advantages

Unlike traditional stereo-to-surround methods that rely on basic Mid/Side processing—which often simply spreads the sides and leaves a muddy or tinny center—this model was trained on native 5.1 and 6-channel mixes segmented into training stems.

Because it learned from authentic multichannel data, the algorithm understands natural sound distribution and channel bleed:

  • Pristine Center Channel: Dry, center-panned dialogue, vocals, and mono SFX are isolated cleanly without the artifacts or frequency bleed typical of simple center extractors.

  • Immersive Stereo Field: Wide vocals, atmospheric SFX, and instruments are beautifully and accurately distributed across the front and rear stereo pairs, creating a truly deep spatial environment.

⚠️ Limitations

A true stereo signal is required for the algorithm to calculate spatial placement accurately.

  • Mono (1ch): Not supported.

  • Dual-Mono (2ch): Highly discouraged. Processing dual-mono audio will confuse the spatial placement algorithms, resulting in unpredictable positioning and unnatural sound.

🗎 Copy link Use algorithm Demo
  • ‹
  • 1
  • 2
  • 3
  • ›
MVSEP Logo
Contact support
Google Play App Store
Site information

FAQ

Quality Checker

Algorithms

Full API Documentation

Company

Privacy Policy

Terms & Conditions

Refund Policy

Cookie Notice

Extra

Help us translate!

Help us promote!