Qwen3 ForcedAligner — Core AI

Apple Core AI conversion of Qwen3 ForcedAligner, source revision c07281df297b9905d24a508279258cccf987a064. The upstream model and converted weights use Apache License 2.0. The complete checkpoint has approximately 918 million parameters; the 0.6B name describes its text backbone.

This model predicts timestamps for supplied words. It does not recognize or replace a transcript. Requires physical Apple silicon with macOS 27 or iOS 27, public Core AI APIs, and a model-specific host runtime. The graphs do not accept raw audio and text directly: the host performs 16 kHz audio framing, log-mel features, byte-level BPE, embedding lookup, causal-cache management and timestamp postprocessing. Geometry and numerical conventions are in metadata.json.

The convolutional frontend, 24-layer audio Transformer and 28-layer text Transformer use W8A16. The small timestamp classifier and embedding table retain FP16 weights. The text network is causal, but all supplied words are known: there is no autoregressive token-generation loop. Timestamp classes are 80 ms apart. The source's monotonic repair can produce zero-duration word spans.

Short inputs use shared-weight entry points for 512, 1,024 or 1,536 positions. Longer inputs use 512-position prefill with the original absolute rotary positions and complete per-layer key/value history, up to 8,192 combined audio and text positions. Audio is limited to five minutes per alignment call; both limits must be respected. Earlier context is retained across prefill blocks. Each text-stage asset contains four layers and one shared weight payload for all entry points. Query/key head normalization is batched across independent heads while preserving the model's normalization semantics.

Upstream advertises 11 languages, listed above. Other languages can be attempted on a best-effort basis; the list is not a language-conditioning restriction. This conversion does not establish timing accuracy outside that coverage. Hosts should preserve native recognizer timings when they are preferable, and retain the original transcript independently of its alignment units.

Measured performance

Audio is John F. Kennedy's “We choose to go to the Moon” speech. These are alignment times for an already supplied transcript, excluding speech recognition, model loading, specialization and audio-file decoding. Measurements include frontend, tokenization, neural stages and timestamp repair. Each workload has one complete warmup before measured repetitions.

Device Audio Warm alignment median RTFx Repetitions
M3 MacBook Air, 16 GB 20-second excerpt 0.103 s 194.9× 5
M3 MacBook Air, 16 GB Complete 18 min 15 s recording 7.541 s 145.3× 5
iPhone 15 Pro Max 20-second excerpt 0.125 s 160.4× 5
iPhone 15 Pro Max Complete 18 min 15 s recording 9.173 s 119.4× 3

The Mac used macOS 27 build 26A428; the phone used iOS 27 build 24A437 and a Release host runtime. The complete recording is 1,095.3 seconds, divided into 36 supplied transcript segments of at most 35 seconds, retaining every sample. These results do not predict the speed of one long-context alignment call. A separate 238.6-second call with 5,151 valid positions takes 6.91 seconds on Mac, approximately 34.5× real time. Long global attention is materially more expensive.

The Mac full-recording measurements were at fair thermal state. The phone's full-recording measurements were taken after cooling, all at nominal state; the short phone measurement was at fair state. A diagnostic run immediately after phone compilation was slower and is not used in this table.

Every phone repetition matched the Mac's supplied input plan, raw timestamp classes and repaired word times exactly, including all 4,448 classes for the complete recording. The export also matches the previously qualified native implementation on the dense and long-context checks. This establishes numerical agreement, not new independent human word-boundary accuracy.

Sampled peak client footprint was about 0.81 GB for the Mac full-recording run and 1.48 GB for the phone full-recording run. These exclude external driver and compiler allocations and are not total device memory use.

A separate iPhone hardware trace exercised all three dense capacities and cached prefill across a 5,151-position input. All 204 graph calls contained ANE prediction activity, with zero target-process GPU intervals. All 1,622 raw timestamp classes and repaired word times matched the Mac. The complete scope and expected per-stage call counts were checked. Placement does not measure arithmetic-unit utilization or energy use.

Preparation and loading

Source .aimodel assets specialize for the device on first use. These are not precompiled .aimodelc assets. Preparation loads assets and functions sequentially.

Device First observed preparation Subsequent cached preparation
M3 MacBook Air 8 min 27 s 0.139 s
iPhone 15 Pro Max 11 min 15 s 0.081–0.579 s

These are observed cache histories, not controlled empty-system-cache or fresh-install guarantees. Related graph experiments preceded the Mac run; underlying OS cache reuse is opaque. Host tokenizer/embedding initialization and audio I/O are excluded from preparation. Storage, application, OS and device changes can affect these times and may trigger specialization again.

AoT compiler products are omitted: no client load-time benefit has been established for this graph. Download one consistent repository snapshot, validate its metadata, and verify files against the checksums supplied by Hugging Face before loading. This package contains runtime assets only, with no recordings, reference transcripts or diagnostic tensors.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for coder543/qwen3-forced-aligner-coreai

Quantized
(1)
this model

Collection including coder543/qwen3-forced-aligner-coreai