01What does BYOK mean?
Bring Your Own Key. You create accounts at AI providers, generate API keys, and enter them into SpeechKey. SpeechKey uses those keys to connect directly to providers — without a SpeechKey server layer in the consumer audio path. You pay providers directly at their rates. SpeechKey adds no markup.
02Why do I need my own API keys?
Because it's the honest model. If SpeechKey managed the keys for you, SpeechKey would be another company sitting between you and your data — with a markup on every API call, a server processing the data flow, and vendor lock-in. BYOK means you pay OpenAI and iFLYTEK directly at their actual rates. SpeechKey is the app layer on your device. The provider relationship is yours.
03Does SpeechKey include AI usage?
No. 9,99 € is the price for the app itself. AI usage costs go directly to your configured providers — typically $0.30–$9.00 per hour depending on pipeline. SpeechKey does not bill these costs and does not see them.
04Can I use SpeechKey without provider accounts?
No. SpeechKey requires at least one configured STT and LLM provider for the translation function. The app does not perform live translation without valid API keys. First-time setup typically takes 10–20 minutes.
05Is BYOK more private than alternatives?
In the consumer BYOK model, SpeechKey does not operate a server in the audio or transcript processing path. There is no SpeechKey account, no SpeechKey analytics, and no SpeechKey transcript storage. This is a different architecture from subscription apps that route audio through their own servers. Whether this is "more private" for you depends on your threat model and the privacy policies of the providers you use.
06Does BYOK mean no data leaves my device?
No. Live AI interpretation requires audio and text to be processed by the providers you configure. BYOK means SpeechKey does not insert its own server layer into the consumer audio/transcript path. What happens to your data at the provider side is governed by the provider's privacy policy and your account settings with them.
07What do I pay to SpeechKey?
9,99 € one-time in the Apple App Store. No monthly subscription, no annual fee, no upsell. One payment — all features included.
08What do I pay to AI providers?
This depends on the pipeline and your usage intensity. Typical meeting costs (with pauses, normal speech density) are:
- China-Only
- ~$0.30–$0.90/hr
- Standard
- ~$0.60–$1.50/hr
- Pro Live
- ~$3.00–$8.00/hr
- Realtime
- ~$3.00–$9.00/hr
These estimates are illustrative. Actual costs depend on provider pricing, model selection, speech density, and exchange rates. Set usage limits in your provider dashboard.
09How can I reduce provider costs?
- Use the Standard pipeline instead of Pro Live or Realtime
- Use iOS system voice instead of MiniMax/ElevenLabs TTS (free)
- Use GPT-4o-mini instead of GPT-4o flagship for translation
- Use China-Only pipeline with iFLYTEK + DeepSeek — cheapest combination
- Silence detection reduces unnecessary API calls
10Why are meeting costs different from continuous speech estimates?
In real meetings there are pauses, silence, and variable speech density. Provider costs are incurred mostly during active speech processing. Continuous speech estimates (uninterrupted talking) can be 30–100% higher than typical meeting costs. We use ranges rather than false precision numbers.
11Does SpeechKey work in China?
Yes — with the China-Only pipeline. It uses iFLYTEK, DeepSeek, and optionally MiniMax, all on Chinese servers. No VPN required. All other pipelines (Standard, Pro Live, Realtime, Watch-Modus) require access to OpenAI, which is blocked in mainland China.
12Do I need a VPN in China?
Only if you use pipelines that include OpenAI (Standard, Pro Live, Realtime, Watch-Modus). The China-Only pipeline requires no VPN — it uses exclusively providers running on Chinese servers.
13Which pipeline works best in China?
China-Only: iFLYTEK for speech recognition, DeepSeek for translation, optional MiniMax for voice output. Cost: ~$0.30–$0.90/hr. No VPN latency, best regional reliability.
14Why does OpenAI sometimes have high latency in China?
OpenAI endpoints are blocked in mainland China. When using a VPN, all API traffic routes through the VPN exit node, which adds latency. VPN speed, exit node location, and local network conditions strongly affect performance. If latency is an issue, switch to the China-Only pipeline.
15Where are my API keys stored?
API keys are stored on your device using iOS security mechanisms such as the iOS Keychain. SpeechKey does not intentionally transmit API keys to SpeechKey servers. In the consumer BYOK model, there are no SpeechKey servers that would receive them.
16Are transcripts sent to SpeechKey?
No. In the consumer BYOK model, SpeechKey does not operate a transcript storage backend. Transcripts are stored locally on your device. You can export, copy, or delete them at any time.
17Are transcripts sent to providers?
Audio and text are sent to the providers you configure for processing — this is the core function of the app. What the provider does with that data is governed by their privacy policy and your account settings. Review each provider's privacy policy and terms of service.
18Are providers allowed to log data?
This depends on your provider account settings and the applicable terms. Many providers offer API usage terms where data is not used for model training by default — but this varies by provider and account tier. Review the settings in your provider dashboard and enable any available data protection options.
19Can I delete transcripts?
Yes. Transcripts stored locally on your device can be deleted from within the app at any time. Since SpeechKey does not store transcripts in the consumer BYOK model, no server-side deletion is required.
20Can I use SpeechKey for confidential meetings?
This depends on your organisation's compliance requirements and the privacy terms of your configured providers. SpeechKey does not operate its own server in the consumer BYOK audio path. Processing is performed by your providers — their terms, data retention, and compliance certifications apply. For enterprise-grade compliance with specific data residency or processing requirements, a separate enterprise configuration may be required.
21Which provider should I start with?
For most users outside China: iFLYTEK (for Chinese speech recognition) + OpenAI (for translation). This is the Standard pipeline and a good balance of quality and cost. For users primarily in China: iFLYTEK + DeepSeek (China-Only pipeline, no VPN).
22What if my key is rejected as invalid?
- Check: no leading or trailing space when pasting
- Check: key pasted into the correct provider field
- Check: billing is active and credit is available at the provider
- For iFLYTEK: all three fields (APPID, API Key, API Secret) must be correct
- If all correct: revoke the key, generate a new one, and re-enter
23How long does setup take?
Typical first setup: 10–20 minutes depending on existing provider accounts. After setup, switching between pipelines takes seconds. The step-by-step guide at setup.html explains each provider step.
24Can I switch providers later?
Yes. You can add new providers or replace existing keys at any time. Pipeline selection in the app allows switching between configurations mid-session. There is no vendor lock-in.
25Is SpeechKey a certified interpreter?
No. SpeechKey is a professional tool for real-time AI interpretation — not a certified legal or medical interpreter. For legally binding or medical interpretation situations, certified human interpreters are required.
26Can it handle technical terminology?
The Terminology Manager injects your own glossaries into the translation system prompt at runtime — proper nouns, project codes, brand names, and technical terms remain consistent. For highly specialised domains (legal documents, medical records) accuracy may vary. Test your vocabulary in a practice session before important meetings.
27Does it work in noisy factories?
Speech recognition quality decreases with increasing background noise. iFLYTEK's ASR is optimised for Mandarin under challenging conditions, but heavy machinery noise or reverb can significantly reduce accuracy. A directional microphone or headset substantially improves recognition in loud environments.
28What happens when the network is weak?
SpeechKey requires an internet connection for all pipelines. In poor network conditions, API calls may be delayed, time out, or fail — resulting in gaps in translation output. Switch to the China-Only pipeline if you are experiencing VPN latency in China. For interrupted connections: the app shows connection errors and automatically retries when connectivity is restored.
29What about languages other than Chinese?
SpeechKey is optimised for ZH↔EN and ZH↔DE. Other language pairs are technically supported — iFLYTEK recognises multiple languages and GPT-4o translates into virtually any language — but pipelines, latency tuning, and terminology templates are optimised for Chinese. For other language pairs quality may vary.
30Can I use SpeechKey offline?
No. All pipelines require an internet connection for speech recognition and/or translation. SpeechKey is designed for professional environments where Wi-Fi or mobile internet is available.
More questions? Email support@speechkey.ai or visit the help centre.