Oruk is a speech lab. We build models for transcription, emotion and speaking-style analysis, with a hosted API and open research.
Resonance analyzes prerecorded English audio. One request can return the transcript, 15 emotion scores, 16 speaking-style scores and timed segments. Scores are independent model outputs, not calibrated probabilities of a person's inner state.
curl https://speech-api.oruk.ai/v1/audio/analysis \
-H "Authorization: Bearer $ORUK_API_KEY" \
-F file=@call.wav -F model=oruk-resonanceThe file API also has separate transcription, emotion and speaking-style endpoints. The Realtime preview streams transcription in 32 locales and phrase-level emotion over WebSocket. Original Resonance’s full emotion and speaking-style analysis uses the English file endpoints.
Resonance-2 Preview is a separate clip-level emotion and speaking-style API. It returns all 31 continuous scores, six signed axes and selected labels that can be empty. Send 0.1–120 seconds of audio, up to 30 MiB, to /v1/audio/resonance-2 using your existing API key and shared speech allowance. This route does not transcribe or diarize audio. SDK 0.2.10 users call it through ordinary HTTP.
Listen to six recorded examples, including mistakes, and inspect the model/calibration revisions and all 546 predictions from the September 17 acted-emotion diagnostic. Training overlap has not been audited; the diagnostic is not an independent held-out evaluation.
Plans include audio minutes, measured by the second. Standard self-serve plans have a seven-day trial that requires a card, charges $0 today and can be canceled before the trial ends. Promotional offers have their own terms. See current plans and allowances.
Orukeet is a 25-language speech recognizer built from NVIDIA Parakeet TDT 0.6B v3. It has NeMo, ONNX and native inference exports.
Start with the tested local-transcription tutorial, including native Q8 and sherpa-onnx CPU examples, pinned artifacts and recorded outputs. The model card documents the weights, licenses and evaluation limitations. OpenWhispr 1.10.0 includes Orukeet as its recommended local model.
oruk-bench contains the published speech-emotion evaluation toolkit and results. The benchmark's seven-class mapping and multilingual dataset describe its evaluation protocol, not the hosted API's native outputs or language support. Oruk's historical model was evaluated in-distribution; those results do not establish the current API's accuracy on independent audio. Read the results and methodology together.
Research articles · RSS · API scope · Contact