Whistle API
    Preparing search index...

    All public entry points are ESM with TypeScript declarations. They have no import-time side effects. Root is the browser entry; use /node explicitly for Node applications.

    import { createWhistle, WhistleError } from '@tiny-stt/whistle';
    import { createWhistle as createNodeWhistle } from '@tiny-stt/whistle/node';
    import { decodeAudio } from '@tiny-stt/whistle/audio';
    import { startRecording } from '@tiny-stt/whistle/microphone';

    Both factories resolve Promise<Whistle> only when inference is ready.

    Browser option Default Contract
    assetsUrl?: string | URL unset CLI-generated directory; relative to document base; no remote fallback
    modelUrl?: string | URL pinned official URL Mirror of exact locked model; integrity verified
    wasmUrl?: string | URL packaged binary Mirror of exact locked WASM; integrity verified
    workerUrl?: string | URL packaged/asset directory worker Matching browser module worker, usually same-origin
    cache?: 'auto' | 'none' auto SDK-managed IndexedDB persistence
    signal?: AbortSignal unset Initialization only; termination cleans up aborted loading
    onProgress?: (p: LoadProgress) => void unset Observer; throwing does not fail initialization

    assetsUrl conflicts with modelUrl and wasmUrl. HTTP(S) URL resolution is required in the browser. Custom model versions are not accepted by URL override.

    Node option Contract
    assetsPath: string | URL Required local directory; file: URLs supported; relative to cwd at invocation
    signal?: AbortSignal Initialization only
    onProgress?: (p: LoadProgress) => void Same progress type; local reads use fromCache: true

    assetsPath never accepts an HTTP(S) location. Every inventory file is verified before the local loader is imported. File paths are resolved independently of the package installation directory. The package's Node worker is located by module URL.

    type WhistleLanguage = 'en' | 'de' | 'fr' | 'es' | 'it' | 'nl' | 'pl';
    interface WhistleWord {
    text: string;
    start: number;
    end: number;
    probability: number;
    }
    interface WhistleTranscript {
    text: string;
    language: WhistleLanguage | null;
    words?: WhistleWord[];
    timings: {
    firstTokenMs: number;
    decodeTokensPerSecond: number;
    totalMs: number;
    };
    }
    interface TranscribeOptions {
    language?: WhistleLanguage | 'auto';
    keywords?: string[];
    wordTimestamps?: boolean;
    signal?: AbortSignal;
    }
    interface Whistle {
    transcribe(audio: Float32Array, options?: TranscribeOptions): Promise<WhistleTranscript>;
    dispose(): Promise<void>;
    }
    interface LoadProgress {
    stage: 'loading-runtime' | 'loading-model' | 'initializing' | 'ready';
    asset?: 'loader' | 'wasm' | 'model';
    loadedBytes?: number;
    totalBytes?: number;
    fraction?: number;
    fromCache?: boolean;
    }

    PCM is 16 kHz mono, 1–480,000 samples, finite and in [-1,1]. Validation and copying happen at invocation, before queueing. Caller buffers are never transferred. A nonzero typed-array offset is respected. Sampling rate is the caller's responsibility.

    Languages default to auto; timestamps default to false. Keyword entries are trimmed, empty entries removed, and case-sensitive duplicates removed. CR, LF and NUL are rejected. The SDK bounds keyword allocations to 128 unique phrases and 8,192 UTF-8 bytes, including newline separators. The header declares no keyword size guarantee; the engine may still reject overly complex prompts. These are SDK allocation limits, not advertised engine quality guarantees.

    Words map the engine's word field to text; times/probabilities are not synthesized. Silence becomes { text: '', language: null } with words: [] when requested. Metrics must be finite; malformed/unsupported output is an explicit error. Transcript boundary whitespace is trimmed, preserving internal punctuation/content.

    totalMs starts in the ready worker before PCM allocation/copy and ends after normalization. It excludes queue delay, cold loading/reloading, recording, decoding, and main-thread delivery. First-token and decoder throughput values come from the engine; they are not microphone-to-caption metrics.

    Each instance owns one worker and engine. Calls run FIFO and the model stays loaded. A queued abort does not disturb another request. An active abort retires that worker generation and reloads before queued work resumes; old events cannot resolve new requests. Failed inference is not automatically retried. A worker crash is terminal (WORKER_FAILED for subsequent work); create a new instance explicitly. Disposal is terminal, idempotent and settles all active/queued work with DISPOSED.

    Progress is per asset. Locked byte sizes are known for transfers; initialization has no numeric estimate. Cache/local completion is marked fromCache: true. ready occurs once per factory, even if a later abort requires reloading.

    decodeAudio(input: Blob, options?: { signal?: AbortSignal }): Promise<Float32Array>;

    interface RecordingOptions {
    maxDurationSeconds?: number; // > 0 and <= 30; default 30
    echoCancellation?: boolean; // default true
    noiseSuppression?: boolean; // default true
    signal?: AbortSignal;
    }
    interface Recording {
    readonly result: Promise<Float32Array>;
    stop(): Promise<Float32Array>;
    cancel(): void;
    }
    startRecording(options?: RecordingOptions): Promise<Recording>;

    Decoding averages all channels equally and resamples through Web Audio. Resampling filter overshoot is bounded to [-1,1]. File clips exceeding 30 seconds are rejected. No remote decoder is used. Browser decoding/rendering has no cooperative abort API: cancellation rejects promptly, closes the live AudioContext, and discards any late decoder/render output; the bounded offline render disconnects its source when done.

    startRecording resolves after capture starts. Manual stop and automatic stop use the same result; repeated stop() calls return that same promise. Cancellation rejects result with ABORTED until completion, including during final decoding. A late permission grant after abort has all its tracks stopped. Permission prompts themselves cannot be programmatically dismissed. Timer/event delay cannot expand output beyond the explicitly requested sample limit. Capture is completed-clip, not continuous transcription. Always observe recording.result rejection.

    WhistleError extends Error with a stable code, message and optional cause. Errors transported from workers preserve a safe code/message pair, not arbitrary engine objects or stack/cause internals. SDK code does not log audio, keywords or transcripts. The upstream loader may write runtime diagnostics on a WASM failure.

    Code Meaning
    UNSUPPORTED_ENVIRONMENT Required browser capability unavailable
    INVALID_OPTIONS Invalid language, URL combination or option type/value
    INVALID_AUDIO Empty/wrong PCM type, non-finite or out-of-range sample
    AUDIO_TOO_LONG More than 480,000 samples or decoded file over 30 seconds
    AUDIO_DECODE_FAILED Invalid/unsupported encoded audio
    MICROPHONE_DENIED Permission/security denial
    MICROPHONE_UNAVAILABLE No usable device, failed start or capture
    ASSET_NOT_FOUND Required local/self-hosted asset missing/unreadable
    ASSET_DOWNLOAD_FAILED Network/HTTP transfer failed
    ASSET_INTEGRITY_FAILED Locked size/SHA-256 mismatch
    ASSET_INCOMPATIBLE Unsupported/malformed manifest or release mismatch
    MODEL_LOAD_FAILED Engine rejected model or did not load speech capability
    INFERENCE_FAILED Negative engine status or runtime inference failure
    INVALID_ENGINE_OUTPUT Truncated, malformed or unsupported engine output
    WORKER_FAILED Worker startup/crash/protocol failure
    OUT_OF_MEMORY Allocation failure or allocation RangeError
    ABORTED Caller cancellation
    DISPOSED Operation on a disposed instance or work rejected by disposal