Modul master Level 3 VibeKoding: Aplikasi Desktop Suara-ke-Teks dengan Electron.Modul master Level 3 VibeKoding: Aplikasi Desktop Suara-ke-Teks dengan Electron.
In this tutorial, we will complete a full closed loop: build a speech-to-text desktop app from scratch with Electron, support both cloud API and local model recognition modes, and finally package it into a real desktop application that can be installed and run on Windows, macOS, and Linux.In this tutorial, we will complete a full closed loop: build a speech-to-text desktop app from scratch with Electron, support both cloud API and local model recognition modes, and finally package it into a real desktop application that can be installed and run on Windows, macOS, and Linux.
For this tutorial, you should at least have:For this tutorial, you should at least have:
Apps you use every day, such as VS Code, Slack, Discord, and Notion, have one thing in common: they are all desktop applications built with Electron.Apps you use every day, such as VS Code, Slack, Discord, and Notion, have one thing in common: they are all desktop applications built with Electron.
Electron is an open-source framework that lets you use HTML + CSS + JavaScript (the same stack used for web pages) to build desktop apps that run across Windows, macOS, and Linux. Its principle is simple: package Chromium and Node.js together, and your web page becomes a standalone desktop app.Electron is an open-source framework that lets you use HTML + CSS + JavaScript (the same stack used for web pages) to build desktop apps that run across Windows, macOS, and Linux. Its principle is simple: package Chromium and Node.js together, and your web page becomes a standalone desktop app.
One-sentence understanding: Electron = an "invisible Chrome browser" + Node.js system capabilities.One-sentence understanding: Electron = an "invisible Chrome browser" + Node.js system capabilities.
๐ผ๏ธ placeholder: A diagram showing the Electron architecture: Chromium (for UI rendering) + Node.js (for system access) = desktop applicationplaceholder: A diagram showing the Electron architecture: Chromium (for UI rendering) + Node.js (for system access) = desktop application
An Electron app consists of two process types. Understanding them is the key to development:An Electron app consists of two process types. Understanding them is the key to development:
Main ProcessMain Process
Renderer ProcessRenderer Process
Preload ScriptPreload Script
contextBridge to safely expose selected APIs to the renderer processUses contextBridge to safely expose selected APIs to the renderer processThey communicate through IPC (Inter-Process Communication), like making a phone call: the renderer says "I want to start recording," and the main process receives that request and calls the system microphone.They communicate through IPC (Inter-Process Communication), like making a phone call: the renderer says "I want to start recording," and the main process receives that request and calls the system microphone.
๐ผ๏ธ placeholder: An Electron process architecture diagram showing Main Process, Renderer Process, and Preload Script, plus IPC communication between themplaceholder: An Electron process architecture diagram showing Main Process, Renderer Process, and Preload Script, plus IPC communication between them
In this tutorial, we will build a Speech-to-Text desktop app. Its functionality is straightforward:In this tutorial, we will build a Speech-to-Text desktop app. Its functionality is straightforward:
Two recognition modes are available:Two recognition modes are available:
| Comparison Dimension | Cloud API Mode | Local Model Mode |
|---|---|---|
| Representative Solution | OpenAI Whisper API | whisper.cpp |
| Internet Required | Yes | No |
| Recognition Speed | Depends on network | Depends on hardware (very fast on Apple Silicon) |
| Chinese Recognition Quality | Excellent | Excellent (large-v3 model) |
| Cost | $0.006/minute | Free |
| Model Size | No download required | tiny model 75MB, large model 3GB |
| Best For | Fast onboarding, lightweight usage | Privacy-focused, offline usage, long-term high-frequency usage |
๐ผ๏ธ placeholder: An app preview showing the speech-to-text UI: recording button and waveform animation at top, recognized text below, and a mode toggle in the top-right cornerplaceholder: An app preview showing the speech-to-text UI: recording button and waveform animation at top, recognized text below, and a mode toggle in the top-right corner
If you have searched for "Electron speech recognition," you may have seen recommendations to use the browser's built-in Web Speech API. Please note: this does not work in Electron.If you have searched for "Electron speech recognition," you may have seen recommendations to use the browser's built-in Web Speech API. Please note: this does not work in Electron.
Google has discontinued speech API support for non-Chrome/Edge browser shells. Electron is Chromium-based, but it is not Chrome itself, so window.SpeechRecognition will fail directly.Google has discontinued speech API support for non-Chrome/Edge browser shells. Electron is Chromium-based, but it is not Chrome itself, so window.SpeechRecognition will fail directly.
That is why we need independent solutions such as OpenAI Whisper API or whisper.cpp.That is why we need independent solutions such as OpenAI Whisper API or whisper.cpp.
We will complete the full flow in the following steps:We will complete the full flow in the following steps:
Open your AI coding assistant and enter this prompt:Open your AI coding assistant and enter this prompt:
CODE Please help me create a new Electron project with Electron Forge using the Vite template. The project name is voice-to-text. Please run: npx create-electron-app voice-to-text --template=vite After creation, enter the project directory and install dependencies.
Electron Forge is the official Electron-recommended scaffolding tool. It helps with project initialization, packaging, distribution, and other tedious setup tasks.Electron Forge is the official Electron-recommended scaffolding tool. It helps with project initialization, packaging, distribution, and other tedious setup tasks.
After creation, the project structure is roughly:After creation, the project structure is roughly:
text voice-to-text/ โโโ src/ โ โโโ main.js # Main process entry โ โโโ preload.js # Preload script (bridge) โ โโโ renderer.js # Renderer process entry โ โโโ index.html # App HTML page โโโ forge.config.js # Electron Forge config โโโ vite.main.config.mjs # Main process Vite config โโโ vite.preload.config.mjs # Preload script Vite config โโโ vite.renderer.config.mjs # Renderer process Vite config โโโ package.json
Ask AI to start the development server:Ask AI to start the development server:
CODE Please help me start the Electron development server by running npm start
After a few seconds, a desktop window appears. This is your Electron app. Even though it only shows a default welcome page now, it is already a real desktop program.After a few seconds, a desktop window appears. This is your Electron app. Even though it only shows a default welcome page now, it is already a real desktop program.
๐ผ๏ธ placeholder: Screenshot of first Electron app startup with the default welcome pageplaceholder: Screenshot of first Electron app startup with the default welcome page
Before implementing speech features, we need to understand Electron's most important concept: IPC (Inter-Process Communication).Before implementing speech features, we need to understand Electron's most important concept: IPC (Inter-Process Communication).
Because the renderer process (UI) and main process (system capabilities) are isolated, they must use IPC "phone calls" to collaborate:Because the renderer process (UI) and main process (system capabilities) are isolated, they must use IPC "phone calls" to collaborate:
text Renderer process (UI) Main process (system) โ โ โโโ "I want to start recording" โโโโโโโโโโโ โ โ โโโ Call microphone โ โโโ Process audio โ โโโโโ "Here is the result" โโโโโโโโโโโโโโ โ โ โโโ Display text in UI โ
In code, this communication is bridged via preload.js:In code, this communication is bridged via preload.js:
javascript // preload.js - safely expose APIs to renderer process const { contextBridge, ipcRenderer } = require('electron') contextBridge.exposeInMainWorld('electronAPI', { // Renderer -> Main sendAudio: (audioData) => ipcRenderer.invoke('transcribe-audio', audioData), // Main -> Renderer onResult: (callback) => ipcRenderer.on('transcription-result', callback) })
javascript // main.js - main process listens for messages const { ipcMain } = require('electron') ipcMain.handle('transcribe-audio', async (event, audioData) => { // Call Whisper API or whisper.cpp here const text = await transcribe(audioData) return text })
๐ผ๏ธ placeholder: IPC flow diagram showing message transfer from Renderer -> Preload -> Mainplaceholder: IPC flow diagram showing message transfer from Renderer -> Preload -> Main
The browser (which is the Electron renderer process) provides navigator.mediaDevices.getUserMedia to access the microphone. Ask AI to help implement recording:The browser (which is the Electron renderer process) provides navigator.mediaDevices.getUserMedia to access the microphone. Ask AI to help implement recording:
CODE Please help me modify src/index.html and src/renderer.js to implement: UI: 1. A large circular "Start Recording" button, which turns into a red "Stop Recording" button when clicked 2. Show a simple pulse animation while recording 3. A text display area below for recognition results 4. Two buttons at the bottom: "Copy Text" and "Clear" 5. A settings icon at top-right to switch recognition mode (cloud/local) Recording logic (in renderer.js): 1. On button click, request microphone access via navigator.mediaDevices.getUserMedia 2. Use MediaRecorder to record audio in webm format 3. After stopping, convert audio Blob to ArrayBuffer 4. Send it to main process via window.electronAPI.sendAudio 5. Wait for recognition result from main process and display it
Core recording code:Core recording code:
javascript // renderer.js let mediaRecorder = null let audioChunks = [] async function startRecording() { const stream = await navigator.mediaDevices.getUserMedia({ audio: { channelCount: 1, sampleRate: 16000, echoCancellation: true, noiseSuppression: true } }) mediaRecorder = new MediaRecorder(stream, { mimeType: 'audio/webm;codecs=opus' }) audioChunks = [] mediaRecorder.ondataavailable = (e) => audioChunks.push(e.data) mediaRecorder.onstop = async () => { const audioBlob = new Blob(audioChunks, { type: 'audio/webm' }) const arrayBuffer = await audioBlob.arrayBuffer() // Send to main process for transcription const result = await window.electronAPI.sendAudio(arrayBuffer) document.getElementById('result').textContent = result } mediaRecorder.start() }
๐ผ๏ธ placeholder: Screenshot of recording UI with red recording state button and pulse animation, plus text result area belowplaceholder: Screenshot of recording UI with red recording state button and pulse animation, plus text result area below
Electron blocks permission requests by default. We need to explicitly allow microphone access in the main process:Electron blocks permission requests by default. We need to explicitly allow microphone access in the main process:
CODE Please help me add microphone permission handling in main.js: 1. Use session.defaultSession.setPermissionRequestHandler to handle permission requests 2. Auto-allow when request type is 'media' 3. For macOS, ensure microphone usage description is declared in package.json or entitlements
javascript // Add to main.js const { session } = require('electron') session.defaultSession.setPermissionRequestHandler( (webContents, permission, callback) => { if (permission === 'media') { callback(true) } else { callback(false) } } )
> Note for macOS users: macOS will show a system-level microphone permission dialog. This is normal. Click "Allow."> Note for macOS users: macOS will show a system-level microphone permission dialog. This is normal. Click "Allow."
This is the simplest option. You only need an API key and a few lines of code.This is the simplest option. You only need an API key and a few lines of code.
sk-) and store it safelyCopy the generated key (starts with sk-) and store it safely> Cost reference: Whisper API costs $0.006/minute. That means recognizing 1 hour of audio only costs $0.36, which is very affordable.> Cost reference: Whisper API costs $0.006/minute. That means recognizing 1 hour of audio only costs $0.36, which is very affordable.
Ask AI to implement speech recognition in the main process:Ask AI to implement speech recognition in the main process:
CODE Please help me implement OpenAI Whisper API in main.js: 1. Install node-fetch (if needed) or use built-in fetch in Node.js 2. Create transcribeWithWhisper function that accepts audio ArrayBuffer 3. Convert ArrayBuffer to Blob/File and build FormData 4. Call https://api.openai.com/v1/audio/transcriptions 5. Use model whisper-1 and set language to zh (Chinese) 6. Return the recognized text 7. Read API key from environment variables or config file
Core code:Core code:
javascript // main.js async function transcribeWithWhisper(audioBuffer, apiKey) { const blob = new Blob([audioBuffer], { type: 'audio/webm' }) const formData = new FormData() formData.append('file', blob, 'audio.webm') formData.append('model', 'whisper-1') formData.append('language', 'zh') const response = await fetch( 'https://api.openai.com/v1/audio/transcriptions', { method: 'POST', headers: { Authorization: `Bearer ${apiKey}` }, body: formData } ) const data = await response.json() return data.text }
๐ผ๏ธ placeholder: Running app screenshot showing recognized Chinese speech returned by Whisper APIplaceholder: Running app screenshot showing recognized Chinese speech returned by Whisper API
Ask AI to add a simple settings panel in the renderer process to input API key and switch recognition mode:Ask AI to add a simple settings panel in the renderer process to input API key and switch recognition mode:
CODE Please help me add a settings panel in index.html: 1. Add a gear icon in the top-right corner; click to expand settings panel 2. The panel includes: - Recognition mode switch (Cloud API / Local model) - API Key input (only visible in cloud mode) - Language dropdown (Chinese / English / Auto detect) 3. Save settings to localStorage 4. Close panel when clicking outside
๐ผ๏ธ placeholder: Screenshot of expanded settings panel showing mode switch and API key inputplaceholder: Screenshot of expanded settings panel showing mode switch and API key input
If you do not want to rely on cloud APIs, or if you need offline usage, whisper.cpp is the best choice. It is a C++ port of the OpenAI Whisper model and runs fully locally without internet.If you do not want to rely on cloud APIs, or if you need offline usage, whisper.cpp is the best choice. It is a C++ port of the OpenAI Whisper model and runs fully locally without internet.
Ask AI to install and configure:Ask AI to install and configure:
CODE Please help me install nodejs-whisper in the project: npm install nodejs-whisper After installation, please help me download the whisper tiny model (small size, fast for testing). nodejs-whisper will handle model download automatically.
> Model selection guide:> Model selection guide:
> * tiny (75MB): fastest, good for testing and lightweight usage, average accuracy> * tiny (75MB): fastest, good for testing and lightweight usage, average accuracy
> * base (142MB): balance between speed and accuracy> * base (142MB): balance between speed and accuracy
> * small (466MB): clearly better Chinese recognition quality> * small (466MB): clearly better Chinese recognition quality
> * large-v3-turbo (1.5GB): recommended; 5-8x faster than large, with only 1-2% lower accuracy> * large-v3-turbo (1.5GB): recommended; 5-8x faster than large, with only 1-2% lower accuracy
> * large-v3 (3GB): highest accuracy, but slower and needs better hardware> * large-v3 (3GB): highest accuracy, but slower and needs better hardware
Ask AI to implement local recognition:Ask AI to implement local recognition:
CODE Please help me add whisper.cpp local recognition in main.js: 1. Import nodejs-whisper 2. Create transcribeWithLocal function 3. Accept audio ArrayBuffer and save it as a temporary WAV file first (16kHz mono) 4. Call nodejs-whisper for recognition 5. Return recognized text 6. Delete temporary file after recognition
Core code:Core code:
javascript // main.js const { nodewhisper } = require('nodejs-whisper') const path = require('path') const fs = require('fs') const os = require('os') async function transcribeWithLocal(audioBuffer) { // Save as temp file const tempPath = path.join(os.tmpdir(), `recording-${Date.now()}.wav`) fs.writeFileSync(tempPath, Buffer.from(audioBuffer)) try { const result = await nodewhisper(tempPath, { modelName: 'base', autoDownloadModelName: 'base', whisperOptions: { language: 'zh', word_timestamps: true } }) return result.map(r => r.speech).join('') } finally { // Clean up temp file fs.unlinkSync(tempPath) } }
๐ผ๏ธ placeholder: Screenshot of local model recognition working offline with Chinese speech inputplaceholder: Screenshot of local model recognition working offline with Chinese speech input
If you are using an M1/M2/M3/M4 Mac, whisper.cpp can automatically use Metal GPU acceleration and Apple Neural Engine. Recognition can run faster than real-time, which means 1 minute of audio may only take a few seconds to process.If you are using an M1/M2/M3/M4 Mac, whisper.cpp can automatically use Metal GPU acceleration and Apple Neural Engine. Recognition can run faster than real-time, which means 1 minute of audio may only take a few seconds to process.
For NVIDIA GPU users, whisper.cpp also supports CUDA acceleration, which provides strong performance too.For NVIDIA GPU users, whisper.cpp also supports CUDA acceleration, which provides strong performance too.
After development is complete, we need to package the app into distributable installers.After development is complete, we need to package the app into distributable installers.
Electron Forge is already included in our project, so packaging is simple:Electron Forge is already included in our project, so packaging is simple:
CODE Please help me run the Electron Forge packaging command: npx electron-forge make
This command automatically generates installers for your current operating system:This command automatically generates installers for your current operating system:
.dmg installer image and .zip archivemacOS: .dmg installer image and .zip archive.exe installer (Squirrel format)Windows: .exe installer (Squirrel format).deb (Debian/Ubuntu) and .rpm (Fedora) packagesLinux: .deb (Debian/Ubuntu) and .rpm (Fedora) packagesBuild outputs are in the out/make/ directory.Build outputs are in the out/make/ directory.
๐ผ๏ธ placeholder: Screenshot of files in out/make directory showing generated .dmg or .exe installersplaceholder: Screenshot of files in out/make directory showing generated .dmg or .exe installers
One "pain point" of Electron apps is large package size (because Chromium is bundled). Optimization suggestions:One "pain point" of Electron apps is large package size (because Chromium is bundled). Optimization suggestions:
dependencies are bundled, and keep dev dependencies in devDependenciesEnsure only packages in dependencies are bundled, and keep dev dependencies in devDependencies| Configuration | Estimated Size |
|---|---|
| Pure Electron app (no model) | ~150-200 MB |
| + whisper tiny model | ~250 MB |
| + whisper large-v3-turbo model | ~1.7 GB |
macOS:macOS:
NSMicrophoneUsageDescription in Info.plistMicrophone permissions must declare NSMicrophoneUsageDescription in Info.plistWindows:Windows:
Linux:Linux:
.deb and .AppImage formatsRecommended to provide both .deb and .AppImage formats> Tip: For personal projects or small-scale distribution, you can temporarily skip code signing and directly share packaged files with friends.> Tip: For personal projects or small-scale distribution, you can temporarily skip code signing and directly share packaged files with friends.
Congratulations! You have built a cross-platform speech-to-text desktop app from scratch. Let's recap what we did:Congratulations! You have built a cross-platform speech-to-text desktop app from scratch. Let's recap what we did:
What makes Electron powerful is that you can build desktop apps at the level of VS Code or Slack using a web-tech stack. And with mature AI speech recognition, a feature like speech-to-text, once requiring a specialized team, can now be built by one person.What makes Electron powerful is that you can build desktop apps at the level of VS Code or Slack using a web-tech stack. And with mature AI speech recognition, a feature like speech-to-text, once requiring a specialized team, can now be built by one person.
Advanced directions:Advanced directions:
Let your voice, and let code record everything for you.Let your voice, and let code record everything for you.