Whisper

Whisper

OpenAI
+
+

Related Products

  • Google Cloud Speech-to-Text
    375 Ratings
    Visit Website
  • QEval
    30 Ratings
    Visit Website
  • 4K Video Downloader
    11,518 Ratings
    Visit Website
  • LALAL.AI
    4,805 Ratings
    Visit Website
  • Google AI Studio
    11 Ratings
    Visit Website
  • Screencapt
    122 Ratings
    Visit Website
  • Crowdin
    867 Ratings
    Visit Website
  • LTX
    141 Ratings
    Visit Website
  • TelemetryTV
    275 Ratings
    Visit Website
  • NaviPlan
    Visit Website

About

Baidu’s speech technology provides developers with such industry-leading capabilities as speech-to-text,text-to-speech, and speech wake-up. Combining with the NLP technology, it is applicable for several scenarios, including speech input, speech search, video subtitle, audio content analysis, calling center, book broadcasting, news broadcasting, and order broadcasting. It can convert a speech with a duration of fewer than 60 seconds to characters. It is applicable for mobile speech input, intelligent speech interaction, speech commands, and speech search. It can convert the audio stream into characters and return each sentence's start and end times. It is applicable for such scenarios as long-sentence speech input, audio and video subtitles, and meeting records. It can convert the audio files uploaded in batches into characters and return the recognition results within 12 hours. It is applicable for such scenarios as record quality check, and audio content analysis.

About

We’ve trained and are open-sourcing a neural net called Whisper that approaches human-level robustness and accuracy in English speech recognition. Whisper is an automatic speech recognition (ASR) system trained on 680,000 hours of multilingual and multitask supervised data collected from the web. We show that the use of such a large and diverse dataset leads to improved robustness to accents, background noise, and technical language. Moreover, it enables transcription in multiple languages, as well as translation from those languages into English. We are open-sourcing models and inference code to serve as a foundation for building useful applications and for further research on robust speech processing. The Whisper architecture is a simple end-to-end approach, implemented as an encoder-decoder Transformer. Input audio is split into 30-second chunks, converted into a log-Mel spectrogram, and then passed into an encoder.

Platforms Supported

Windows
Mac
Linux
Cloud
On-Premises
iPhone
iPad
Android
Chromebook

Platforms Supported

Windows
Mac
Linux
Cloud
On-Premises
iPhone
iPad
Android
Chromebook

Audience

Companies and anyone looking for a tool to convert speech and audio to written text and subtitles

Audience

Anyone looking for a tool to recognize speech automatically and improve text transcription

Support

Phone Support
24/7 Live Support
Online

Support

Phone Support
24/7 Live Support
Online

API

Offers API

API

Offers API

Screenshots and Videos

Screenshots and Videos

Pricing

No information available.
Free Version
Free Trial

Pricing

No information available.
Free Version
Free Trial

Reviews/Ratings

Overall 0.0 / 5
ease 0.0 / 5
features 0.0 / 5
design 0.0 / 5
support 0.0 / 5

This software hasn't been reviewed yet. Be the first to provide a review:

Review this Software

Reviews/Ratings

Overall 0.0 / 5
ease 0.0 / 5
features 0.0 / 5
design 0.0 / 5
support 0.0 / 5

This software hasn't been reviewed yet. Be the first to provide a review:

Review this Software

Training

Documentation
Webinars
Live Online
In Person

Training

Documentation
Webinars
Live Online
In Person

Company Information

Baidu
Founded: 2000
China
intl.cloud.baidu.com/product/speech.html

Company Information

OpenAI
United States
openai.com/blog/whisper/

Alternatives

Alternatives

Transcribe

Transcribe

Wreally

Categories

Categories

Integrations

AI Sparks Studio
Android
Baseten
Blink
Fuser
Hyprnote
Kuku
LastMile AI
LazyTyper
NoteVocal
ReByte
Simplismart
Thinkbuddy
Tila
Undrstnd
Unremot
Utterly Voice
VESSL AI
Waveloom
Zo

Integrations

AI Sparks Studio
Android
Baseten
Blink
Fuser
Hyprnote
Kuku
LastMile AI
LazyTyper
NoteVocal
ReByte
Simplismart
Thinkbuddy
Tila
Undrstnd
Unremot
Utterly Voice
VESSL AI
Waveloom
Zo
Claim Baidu AI Cloud Speech-to-Text and update features and information
Claim Baidu AI Cloud Speech-to-Text and update features and information
Claim Whisper and update features and information
Claim Whisper and update features and information