Join/Login
Business Software
Open Source Software
For Vendors
Blog
About
More

For Vendors Add a Product Join Login

Business Software

Open Source Software

SourceForge Podcast

Resources

Articles
Case Studies
Blog

Help
Create
Join
Login

Home
Compare Business Software
Artificial Intelligence
Speech Recognition Software

Best Speech Recognition Software - Page 2

View:

Open Source Commercial

Clear All Filters

Speech Recognition Features

Automatic Transcription 62
Voice Recognition 57
Multi-Languages 55
Speech-to-Text Analysis 52
More...
Audio Capture 49
Specialty Vocabularies 25
Continuous Speech 24
Call Analysis 19
Concatenated Speech 19
Customizable Macros 19
Automatic Form Fill 10
Variable Frequency 10

Deployment

Cloud 91
Windows 34
iPad 27
iPhone 27
Android 24
On-Premises 13
Mac 11
Linux 7
Chromebook 2

Compare the Top Speech Recognition Software as of December 2025 - Page 2

Sort By:

Sponsored

Speech Recognition Clear Filters

1

FirstLanguage

FirstLanguage

Our Natural Language Processing(NLP) APIs provide best-in-class accuracy at an affordable rate and cover all aspects of NLP under a single roof. Save weeks of time training and creating language models. Take advantage of our best-in-class APIs to kickstart your app development. We provide the building blocks to create your own apps effectively like chatbots, sentiment analysis, etc. Text classification on multiple domains and 100+ languages. Perform effective sentiment analysis. We grow when your business does. So we have put together simple pricing that allows you to easily scale your business when it needs to evolve. Perfect for individual developers who are creating apps or building proof of concepts. Head to the Dashboard and get your API Key. Place this in the header of all your API calls. Use our SDK in your preferred language to start coding. Or you can refer to the auto-generated code blocks provided in 18 programming languages.

Starting Price: $150 per month

View Software
2

Picovoice

Picovoice

Picovoice is the first and only ubiquitous on-device voice AI platform. Picovoice offers speech-to-text, voice search, wake word, Speech-to-Intent (intent detection) and voice activity detection engines. Its stack can run on anything from embedded devices to web browsers, providing an immersive experience not achievable by any Big Tech.

Starting Price: Free

View Software
3

Work by Speech

Mikołaj Magowski

Work by Speech is the first program in the world that allows efficient work on a computer by speech without needing a keyboard and mouse. Work by Speech Features: - Efficient work on a computer by speech alone - Quiet speaking support - Application switching and opening by speech - Built-in voice commands for the most common actions - Custom voice commands management - Macro recording and editing - Separate dictation mode - Fast and repeatable mouse control by speech with support for all mouse actions - Customizable mousegrid that can be moved by speech - Automatic mousegrid optimization for every used application - Very low processor and memory usage - Works with any microphone under Windows 10 and 11 - Available for the English language only - Free updates

Starting Price: Free

View Software
4

SpeechPulse

AV BEAM

SpeechPulse uses your computer’s microphone for real-time speech recognition. It can type into your favorite apps, including text editors, web browsers, and office applications. SpeechPulse works fully offline and doesn’t require any internet connectivity. It supports speech recognition in multiple languages, including English, French, Spanish, Italian, German, Japanese, Chinese, and Russian (a total of 100 languages). SpeechPulse supports both auto punctuation and manual punctuation for the English language. It supports auto punctuation for all other languages. SpeechPulse can also generate subtitles for your audio and video files with accurate timestamps. It supports SRT and VTT subtitle formats. You can also customize the width of a subtitle line to include only a limited number of characters. SpeechPulse has a one-time payment. You can pay for the product once and use it forever.

Starting Price: $59.95/one-time payment

View Software
5

Yandex SpeechKit

Yandex

Speech technologies based on machine learning to create voice assistants, automate call centers, monitor service quality, and perform other tasks. Leverage the advanced technology behind the wildly successful Alice voice assistant, now ready for use in your business. In a fraction of a second, SpeechKit accurately recognizes speech, allowing our clients' voice assistants to communicate quickly and easily. Choose the right version for you, the full version creates a smart voice assistant while the adaptive version gives your brand a unique voice in just a month. A solution for the most demanding customers who need to control speech processing and synthesis within their own infrastructure. SpeechKit’s ML models can now be deployed to your infrastructure. We offer both hybrid options and 100% on-premise deployments for sensitive traffic. The service can recognize audio in MP3, LPCM, and OggOpus formats.

Starting Price: $0.000020 per unit

View Software
6

Gladia

Gladia

Gladia is an advanced audio transcription and intelligence platform delivered via a unified API that supports both asynchronous (pre-recorded) and real-time streaming transcription, enabling developers to convert speech to text in over 100 languages with features like word-level timestamps, language detection, code-switching, speaker diarization, translation, summarization, custom vocabulary, and entity extraction. Its real-time engine achieves latencies under 300 ms while maintaining high accuracy, and it offers “partials” (intermediate transcripts) to improve responsiveness in live settings. The platform’s asynchronous API is powered by a proprietary Whisper-Zero model optimized for enterprise audio, and it lets clients apply add-ons such as enhanced punctuation, name consistency, custom metadata tagging, and export to subtitle formats (SRT, VTT).

Starting Price: Free

View Software
7

Go Transcribe

Go Transcribe

Sign up for a free account. Upload your audio/video files straight onto our web based transcription platform. Statistics prove that including subtitles results in your videos standing out. Additionally, over 80% of media played on social media platforms are played in mute, so including subtitles can easily capture your viewer’s interest! By including subtitles in your media, your viewers will get your point effortlessly. For example, if you are asking your viewers to donate to a meaningful charity. If you include subtitles, the chances of getting donations will increase because you will be understood, this also goes if you are asking for sales! Additionally, it helps people who have problems with hearing. These are a few reasons why adding subtitles is a massive help for your business. But if you didn’t know, creating subtitles isn’t easy. It is prolonged and expensive! You don’t need to worry, though.

Starting Price: $10.80 one-time payment

View Software
8

Calldrip

Calldrip

What is Calldrip and why should my sales organization use it? For more than 10 years Calldrip has been dedicated to helping businesses respond immediately to new inquiries. We've leveraged this experience to develop our suite of sales automation tools and have now deployed this technology to thousands of customers worldwide. By triggering a phone call between your sales team and your prospect while they're still on your website, were able to increase conversations by as much as 900%. The privately held, fast-growing company is based in Salt Lake City, UT. In todays Google Micro Moments world, business must engage with the prospect FAST. Calldrip ensures instant engagement and spotlights potential problems in the sales processes.

Starting Price: $99.00/month/user

View Software
9

BigHand Dictation and Speech Recognition

BigHand

Boost productivity and profitability by empowering your teams to spend less time transcribing, and more time on higher-priority work. Enable accurate dictation that’s not only fast to complete, but incredibly straightforward to manage with configurable workflows. Staff can record simply using their voice via desktop, mobile or tablet, and easily share, prioritize and track files.

View Software
10

LumenVox Automatic Speech Recognition (ASR)

LumenVox

Transforming customer engagement with AI-powered voice recognition and voice authentication technology. Our flexible voice-enabled technology allows you to create a solution that meets all of your customers' demands, affordably and reliably. We do one thing, and we do it well. And that's voice enablement for your apps. Finally, deliver great voice automation and interactions. Whether it's short, simple commands or conversational questions, LumenVox ASR and TTS are accurate and affordable, helping you improve efficiency on both sides of the phone line. You will never repeat yourself. Recognize multiple dialects from a single global language model to serve all your customers. We give you maximum flexibility from a capabilities, implementation and monetization perspective. If you can think it, you can build it with LumenVox

View Software
11

Phonexia Speech Platform

Phonexia

Phonexia offers a comprehensive portfolio of cutting-edge speech recognition and voice biometrics technologies ready to meet any commercial and governmental scenarios. Powered by the latest advancements in artificial intelligence, acoustics, phonetics, and voice biometrics science, Phonexia products are extremely accurate, fast, and scalable. Phonexia’s AI-powered solutions let you build voicebots, verify a speaker’s identity based on voice biometrics, transcribe speech to text, and search for speakers and context in large amounts of audio. Secure access to your clients’ data conveniently with voice biometric authentication and detect fraud attempts natively. Phonexia offers a comprehensive portfolio of cutting-edge speech recognition and voice biometrics technologies ready to meet any commercial and governmental scenarios. Powered by the latest advancements in artificial intelligence, acoustics, phonetics, and voice biometrics science.

View Software
12

TranscribeMe

TranscribeMe

The way we think about data is changing; and now, more than ever, industry leaders are counting on reliable, highly accurate transcription and data annotation for their business. Our proprietary task distribution and workforce management platform has been built with the industry’s best information security protocols and processes to ensure that your data is encrypted and securely maintained. We offer workflows compliant with HIPAA and GDPR protocols, and all of our services can be customized; including geofencing the workforce to specific locations. The technology and workflows we have built enable us to deliver the highest quality data consistently and at low prices. Successful artificial intelligence and machine learning models require data that is relevant to your use case. As experts in curating large groups of workers, we can deliver the best data for a variety of use cases that include creating contact center interactions, images, review and survey data, and much more.

Starting Price: $0.79 per minute

View Software
13

WebsiteVoice

WebsiteVoice

Turn all your website articles into high-quality audio in less than 5 minutes and for free. Let your visitors listen to the content of your website in the background while they do other things with our text-to-speech technology and increase the time spent on your website. Accessibility is sometimes forgotten. Empower visitors with visual impairment and reading disabilities to still completely consume your content without the complications of reading. Listening to podcasts and audiobooks has become a growing trend and behavior for people to consume content. Capture a wider audience that would prefer tuning in instead of reading. Thanks to our Automatic Content Recognition technology, you can just drop our snippet on your site and forget about it. We will automatically enable text-to-speech voice for the relevant content. We use Artificial Intelligence and Machine Learning to constantly improve our voice algorithms to make your website text-to-speech as realistic as possible.

Starting Price: $9 per month

View Software
14

Symbl

Symbl.ai

Symbl is an API platform for developers and businesses to rapidly deploy conversational intelligence at scale – on any channel of communication. Our comprehensive suite of APIs unlock proprietary machine learning algorithms that can ingest any form of conversation data to identify actionable insights across domains and channels (voice, email, chat, social) contextually – without the need for any upfront training data, wake words, or custom classifiers. Symbl is democratizing conversational tech to make collaboration effortless at scale. We provide the technology for organizations to deploy at scale our proprietary workplace productivity API so brands can optimize key workflows for knowledge workers or enhance the customer experience. Whether you are a seasoned developer or just starting to explore how to harness employee collaboration to fit your organization’s needs, our API can be customized for your specific applications.

View Software
15

Azure Speaker Recognition

Microsoft

A Speech service feature that verifies and identifies speakers. Enable frictionless, secure customer experiences: Improve the customer experience by streamlining verification processes. Use voice to verify individuals for secure, frictionless customer engagements in a wide range of solutions, from web applications to call centers. Speaker verification can use either passphrases or free-form voice input. Improve the customer experience by streamlining verification processes. Use voice to verify individuals for secure, frictionless customer engagements in a wide range of solutions, from web applications to call centers. Speaker verification can use either passphrases or free-form voice input. Unlock value from scenarios with multiple speakers: Determine a speaker’s identity from within a group of enrolled speakers. Speaker identification enables you to attribute speech to individual speakers, support multiuser voice recognition for personalized interactions, and more.

View Software
16

Voice Pro

LinguaTec

Voice Pro Enterprise has been developed especially for use in enterprises. The recognition is done on the company server and can be accessed from any device (PC, Mac, smartphone, tablet). This ensures that all in-house information remains within the company. No more time-consuming speaker training is necessary, thanks to the speaker-independent recognition technology: Just speak into your device and you will see the transcribed text immediately. Companies finally have a sophisticated and secure speech recognition solution at their disposal. Regardless of whether you need to create a document at your work station, write an email on the move or dictate a sales report on site: Voice Pro Enterprise saves time and helps to make employees more productive. Voice Pro Enterprise results in a noticeable increase in employee efficiency. With Voice Pro Enterprise you dictate on average three times faster than you type. The high recognition accuracy minimizes post-processing.

Starting Price: €149 one-time payment

View Software
17

Deepgram

Deepgram

Deploy accurate speech recognition at scale while continuously improving model performance by labeling data and training from a single console. We deliver state-of-the-art speech recognition and understanding at scale. We do it by providing cutting-edge model training and data-labeling alongside flexible deployment options. Our platform recognizes multiple languages, accents, and words, dynamically tuning to the needs of your business with every training session. The fastest, most accurate, most reliable, most scalable speech transcription, with understanding — rebuilt just for enterprise. We’ve reinvented ASR with 100% deep learning that allows companies to continuously improve accuracy. Stop waiting for the big tech players to improve their software and forcing your developers to manually boost accuracy with keywords in every API call. Start training your speech model and reaping the benefits in weeks, not months or years.

Starting Price: $0

View Software
18

Azure AI Speech

Microsoft

Build voice-enabled apps confidently and quickly with the Speech SDK. Transcribe speech to text with high accuracy, produce natural-sounding text-to-speech voices, translate spoken audio, and use speaker recognition during conversations. Create custom models tailored to your app with Speech studio. Get state-of-the-art speech to text, lifelike text to speech, and award-winning speaker recognition. Your data stays yours, your speech input is not logged during processing. Create custom voices, add specific words to your base vocabulary, or build your own models. Run Speech anywhere, in the cloud or at the edge in containers. Quickly and accurately transcribe audio in more than 92 languages and variants. Gain customer insights with call center transcription, improve experiences with voice-enabled assistants, capture key discussions in meetings and more. Use text to speech to create apps and services that speak conversationally, choosing from more than 215 voices, and 60 languages.

View Software
19

Dragon Legal

Nuance Communications

Dragon Legal is a specialized speech recognition software tailored for legal professionals, offering a legal-specific language model trained on over 400 million words from legal documents. This enables attorneys and legal practitioners to dictate contracts, briefs, and legal citations with up to 99% accuracy, three times faster than typing. The software supports the creation of custom voice commands to automate repetitive tasks and allows for the transcription of pre-recorded audio files, enhancing workflow efficiency. Optimized for Windows 11 and compatible with Windows 10, Dragon Legal v16 also provides accessibility features such as "play that back" audio of dictated text and sophisticated macro commands, accommodating legal professionals with physical or cognitive disabilities. Additionally, it offers integration with Dragon Anywhere Mobile, a cloud-based dictation solution for iOS and Android devices, ensuring productivity on the go.

Starting Price: $799 one-time payment

View Software
20

Voice Finger

Voice Finger

Enables zero computer contact, no need for keyboards and mouses. Rest your hands and use your voice to command the computer. A definitive solution for people with disabilities and/or computer injuries. Some speech recognition software assumes you can type and click for some tasks. Voice Finger was made to do everything by voice. Also for hardcore gamers. For competitive gamers, Voice Finger can hit keys and buttons while the gamer moves and shoots, acting like a third hand. Voice Finger allows complete control of the keyboard, with short commands to navigate the cursor, type, hold and hit keys and buttons. Windows default speech recognition has a lot of lengthy commands like "Press 1", "Press A" and "Press down 30 times". Voice Finger cuts down all commands to a minimum length, like "1", "A" and "Down 30", and you are still able to use the mouse buttons with commands like "click left", "click right" and others, and at the same time hold keys like Control, Shift and Alt.

Starting Price: $9.99 one-time payment

View Software
21

VoxCommando

VoxCommando

VoxCommando is a speech recognition and command utility that lets you take control of your multimedia Home Theatre PC (HTPC). VoxCommando can be run locally, without sacrificing privacy to any cloud-based services. Add voice control to your home automation. Use it as an assistive tool to speed up everyday tasks, reduce your reliance on the keyboard and mouse. VoxCommando is different from other speech recognition applications in that it is extremely customizable. It is designed to work with a wide variety of home automation services and multimedia programs, including user favorites like Kodi and MediaMonkey. It is able to achieve accurate speech recognition because it already knows what media is in your library.

View Software
22

aiOla

aiOla

aiOla is a deep tech Conversational, Voice, and Speech AI lab with an enterprise-level automatic speech recognition (ASR) foundation model, Text-to-speech (TTS) technology and Natural Language Understanding (NLU). It’s designed to help enterprises and developers adapt speech technologies to any process, whether through seamless API integration or an intuitive in-house app. aiOla is revolutionizing enterprise operations with enterprise level Conversational AI. We specialize in speech-to-text and text-to-speech AI that deliver unmatched accuracy (95%), specialized in specific jargon, in any language, accent, vertical, or acoustic environment. From empowering frontline workers with hands-free workflows to enabling voice AI agents with enterprise-grade ASR and TTS, aiOla seamlessly integrates into workflows, internal apps and products.

View Software
23

Txtplay

Txtplay

Txtplay not only makes your video and audio accessible for everyone it also extracts hidden powers in your media: searchable metadata. This means archiving, SEO, compliance become much easier to manage. Upload your media and select your language. Our speech recognition engine will take care of the job and notify you when it's done. You can continue working while our AI is doing the magic. We connect your media to the transcript in our online text editor where you can update, highlight, detect speakers and search through your text, and scroll in your audio or video. We support over 20 formats including: SRT, VTT,.docx. You can fine-tune the export with details like Timecode, Atlas format, speakers, etc. We also have developer-friendly options.

Starting Price: €0.25 per min

View Software
24

Line 21

Line 21

Line 21 provides AI-powered live captions and subtitles, ensuring seamless accessibility for live events, streaming platforms, and digital content. Our hybrid approach combines AI automation with human expertise, delivering high-accuracy captions that adapt to industry-specific terminology, accents, and niche references. By leveraging our AI Proofreader, we enhance real-time captions, reducing errors and making live experiences more inclusive and engaging. Our solution is designed for event organizers, broadcasters, and language service providers who need scalable, cost-effective, and high-quality captions. Traditional human captioning is expensive and non-scalable, while ASR solutions often lack accuracy. Line 21 bridges this gap by offering real-time AI-enhanced captions that integrate seamlessly into event tech and streaming workflows.

Starting Price: $0.09/min

View Software
25

Diktamen

Diktamen

Diktamen is a cloud-based digital dictation and transcription platform designed to streamline voice capture, task management, and workflow automation across professional sectors. The solution enables users to dictate audio from any location, via mobile, desktop, or dedicated devices, and securely transmit that audio for transcription, speech recognition, and task assignment. It supports industry-specific workflows (notably in legal and healthcare), allows integration with existing systems, and features centralized management for submissions, status tracking, and BI reporting with AI-driven forecasting. Clients benefit from cost reduction in dictation infrastructure, efficient transcription turnaround through outsourced partner networks, real-time task routing, and a flexible SaaS deployment model with minimal local installation or maintenance. Diktamen holds ISO 27001 certification and adheres to GDPR for data security and compliance.

View Software
26

SmartAction

SmartAction

SmartAction brings together best-of-breed technologies and services to deliver conversational AI as a fully managed experience. With more than 100 customer deployments, we know a thing or two about automating conversations that drive engagement and resolution. Why trust your CX to anything less? Building and managing a virtual agent has never been so easy — because we do it all for you. From the conversation design, implementation, and continuous optimization, the SmartAction CX team supports you at every step in the conversational AI journey. Because every customer interaction is different, SmartAction tailors its natural language understanding (NLU) system question by question, to achieve the highest accuracy possible. This approach enables our intelligent virtual agents to perform on par and sometimes even better than live agents.

View Software
27

SpokenData

ReplayWell

Let the automatic speech-to-text technology transcribe your data. Or transcribe your data yourself or buy professional transcript. Use our on-line time synchonous editor to surf your data and transcripts. Download transcripts in many formats. Manage your team of transcribers using tags and categories. Help them with transcription by automatic voice-to-text technology. Integrate SpokenData into your application via our REST API. We adapt the voice-to-text on your data domain to maximize the transcript accuracy and lower your labor costs. Enable speech technologies in your applications through integrating SpokenData using our REST API. We are ready to process huge amounts of your data. You get API fitting your needs. Just contact our support team. We customize the voice-to-text on your data and purpose to maximize the transcript accuracy. Suitable for: web/mobile app developers, media monitoring agencies, audio/video archive business.

View Software
28

VoxSigma

Vocapia

The VoxSigma software suite is offered as a Web service via a REST API over HTTPS, always providing customers access to our latest systems thereby quickly benefiting from regular advances and take advantage of additional features offered by the online environment. Our speech-to-text service is available 24/7/365 with failover servers and geographic redundancy. Automatic on-the-fly adaptation allows the user to provide texts related to the audio document being processed, what can be considered topic/domain adaptation. These accompanying texts serve to increase the lexical coverage of the speech-to-text system and to adapt the language model to the specific domain of the audio document with the aim of improving the transcription accuracy.

View Software
29

Trint

Trint

Introducing the easiest way to record, transcribe and share right from your phone! Trint’s mobile app lets you capture the moments that matter, anywhere, anytime. Wired: “Amazing!” Google: “Rocket-fueling innovation!” We understand work doesn’t always happen in an office, so we built the mobile app to give you all the power of Trint’s AI transcription on-the-go. Record live interviews and import files from your phone directly without any clunky equipment. It’s all in the app! Record live conversations. Import audio files into Trint from your other apps. Share transcripts and set editing permissions in-app. Intuitive player to easily follow Trint transcripts. All files saved to your device or to the cloud so never worry about losing a file. Download audio to your device. Drop markers from your Apple Watch while you record. Capture in 28 languages, right from your phone, including English, Spanish, French, Chinese Mandarin, Hindi, etc.

View Software
30

Yactraq

Yactraq

Yactraq is the industry value leader in speech analytics software. Our customers typically realize benefits across two broad functional areas. Marketing teams looking to extend their Voice-of-the-Customer (VoC) capabilities beyond the feedback form and social media now want to mine sales and customer service phone calls as part of their omni-channel capability. Contact Center Quality Management teams typically use speech analytics / audio mining as a way of leveraging AI / Machine Learning to evaluate the performance of their call agents. Yactraq offers customized free trials based on a clients own data so they can experience the value of our software before deciding to buy. Our products are cost-effectively priced to suit the needs of end customers as well as partners in the Business Process Outsourcing (BPO), Contact Center as a Service (CCAS), Voice-of-the-Customer (VoC), CRM Software and Network Service Provider businesses.

View Software

Previous
1
You're on page 2
3
4
Next

Related Categories

Language Learning Medical Transcription Speech Analytics Biometric Authentication Voice Bot Speech to Text Multimodal Models AI Call Center Dictation AI Receptionists Live AI Translation

SourceForge

Open Source Software
Business Software
Add Your Software
Business Software Advertising

Company

About
Team
SourceForge Headquarters
1320 Columbia Street Suite 310
San Diego, CA 92101
+1 (858) 422-6466

Resources

Support / Documentation
Site Status
SourceForge Reviews

Terms Privacy Opt Out Advertise

Thanks for helping keep SourceForge clean.

You seem to have CSS turned off. Please don't fill out this field.

Briefly describe the problem (required):

Upload screenshot of ad (required):

Select a file, or drag & drop file here.

✔

✘

Screenshot instructions:

Click URL instructions:
Right-click on the ad, choose "Copy Link", then paste here →
(This may not be possible with some types of ads)

More information about our ad policies

Ad destination/click URL:

Best Speech Recognition Software - Page 2

Compare the Top Speech Recognition Software as of December 2025 - Page 2

FirstLanguage

Picovoice

Work by Speech

SpeechPulse

Yandex SpeechKit

Gladia

Go Transcribe

Calldrip

BigHand Dictation and Speech Recognition

LumenVox Automatic Speech Recognition (ASR)

Phonexia Speech Platform

TranscribeMe

WebsiteVoice

Symbl

Azure Speaker Recognition

Voice Pro

Deepgram

Azure AI Speech

Dragon Legal

Voice Finger

VoxCommando

aiOla

Txtplay

Line 21

Diktamen

SmartAction

SpokenData

VoxSigma

Trint

Yactraq

Related Categories