CTVG-4B

Confidence-based Temporal Video Grounding

Generate once. Rank, select, or reject.

CTVG-4B is the final 4B-family model from Grounding with Confidence: Controllable Generative Video Temporal Grounding. Given a video and a text query, it generates candidate time intervals and scores each interval with a lightweight confidence head using decoder states from the same generation pass.

Paper on arXiv · Code on GitHub · Base model

This repository contains a complete merged generator and its matched confidence head. Loading only the generator with Transformers does not apply the head or reproduce confidence-based selection. Use the accompanying Grounding-with-Confidence inference code.

Model details

Property Value
Model name CTVG-4B
Model repository Alibaba-VELLDEPTH/CTVG-4B
Task Text-conditioned video temporal grounding with interval-level confidence
Base model MCG-NJU/TimeLens2-4B (Qwen3-VL architecture)
Base revision ddbb6cb944f13ce21e59e85da23c5f356107260e
Training Supervised fine-tuning followed by set-level reinforcement learning
Released checkpoint Final low-learning-rate RL generator, step 200
Confidence head Matching online head after 150 updates
Generator format Merged BF16 SafeTensors; not a LoRA-only adapter
Confidence architecture LayerNorm(2560) → Linear(2560, 256) → GELU → Dropout(0.1) → Linear(256, 1), followed by sigmoid
Confidence features Mean of semantic interval-token decoder states, with stop-gradient during head training

The model separates candidate generation from acceptance. Generation-order NMS at temporal IoU 0.3 runs before confidence thresholding; confidence does not reorder NMS. Threshold changes can reuse the saved candidate scores without another decoding pass. No external verifier or additional confidence tokens are needed at inference.

Getting started

Get the inference code

Clone the code repository, then run the following download and inference commands from its root:

git clone https://github.com/Alibaba-VELLDEPTH/Grounding-with-Confidence.git
cd Grounding-with-Confidence

Download the complete model

Use the Hugging Face CLI in a separate download environment if needed:

hf download Alibaba-VELLDEPTH/CTVG-4B \
  --local-dir models/CTVG-4B

For reproducible experiments, add --revision COMMIT_SHA using the commit shown in this repository's Files and versions tab. Access to a private repository requires login with an authorized account.

Download the whole repository, not just the generator shards. release.json, span-confidence-head.safetensors, and span-head-config.json are required for model/head pairing validation. No separate base-model download is needed for this merged release.

Run confidence-aware inference

From the root of the accompanying Grounding-with-Confidence codebase, use Python 3.10+ and a CUDA GPU:

python -m pip install -e .
python -m pip install -r requirements.txt

python scripts/infer.py \
  --model models/CTVG-4B \
  --input requests.jsonl \
  --output-dir outputs/ctvg-example \
  --threshold 0.5

Create requests.jsonl with one request per line, replacing the example path, query, and duration with your own data:

{"request_id":"example-1","video_path":"videos/example.mp4","query":"A person opens a door.","source_duration_s":30.0}

A relative video_path is resolved against the request JSONL's directory. Add gt_intervals to every request and pass --evaluate to also produce an evaluation report. Add --check-only to validate inputs and artifact hashes without loading the GPU model.

The output predictions.jsonl includes candidate intervals, their span_scores, and selected_intervals. The example cutoff of 0.5 is illustrative, not a calibrated default; select a threshold on separate validation data for your application. Failed requests remain explicit in the output and evaluation denominator.

The released inference configuration uses 2 fps, min_pixels=2048, total_pixels=8388608, a 16,384-token context, and at most 512 output tokens. SDPA is the default attention backend. The full model/head inference path requires the companion scripts rather than a generic text-generation pipeline.

Training data and procedure

Public video-grounding data sources are TimeLens2-93K and OMTG-56K. Obtain the data from the upstream repositories under their respective terms; source videos are not redistributed here.

  • SFT: GT-anchored interval targets train the generator with LoRA. A separate fixed candidate replay and verifier-v2 continuous soft labels train the detached confidence head with soft BCE and ranking supervision. The historical run used 60,405 units and 1,888 optimizer steps.
  • RL: Set-level rewards train the generator, while online overlap supervision updates the detached head after actor updates. The released checkpoint is actor step 200 with 150 online head updates, not the bootstrap SFT head.

The upstream downloads alone do not reconstruct the paper-specific splits, offline replay, or verifier scores. This weight release does not include original training videos, those additional supervision files, optimizer state, or distributed training checkpoints.

Evaluation

Reported results on all 320 OMTG-Bench queries:

Metric Reported value (%)
tF1@0.5 67.39
Temporal IoU 63.24
EtF1 42.64

These are the manuscript's retrospective benchmark results. The operating threshold, 0.18404516577720642, was selected on the test set and must not be treated as an independently calibrated deployment threshold. The default companion evaluator follows the official OMTG metric protocol, including its merge-before-matching behavior; it is distinct from generic interval-set metrics.

Confidence also improves selection from the same fixed post-NMS candidate pool under equal global return budgets. Official query-macro Recall@0.5 (%), with 1,314 candidates across 320 queries:

Global return budget Generation order Token likelihood External verifier Confidence
10% 9.95 9.41 11.00 14.42
25% 26.48 22.37 26.31 31.12
50% 51.34 43.05 47.72 52.82
75% 64.76 61.00 62.45 65.30

The budget is shared across queries, not assigned independently to each query. These tables report prior experiments, not a new inference run performed during upload. GPU end-to-end acceptance of the companion release wrapper remains a separate validation step; CPU artifact and metric checks do not establish GPU execution correctness.

Intended use and limitations

CTVG-4B is intended for research on video moment retrieval, repeated-event localization, and controllable selection or rejection of generated temporal intervals, subject to applicable permissions.

  • Confidence can only select among generated candidates. It cannot recover missing events, repair boundaries, or reverse NMS suppression.
  • Scores do not establish calibrated correctness probabilities. Precision targets and thresholds may not transfer to new datasets or applications.
  • Most of the full-system tF1 improvement comes from NMS. The additional confidence gain after NMS is +0.32 percentage points with a 95% source-video bootstrap interval of [−0.14, 0.76], so that incremental improvement is not statistically established.
  • Reported rejection experiments use synthetic cross-video query mismatches, not general open-world rejection.
  • Do not rely on the model as the sole basis for safety-critical decisions or consequential decisions about people. Review outputs and comply with privacy and video-use permissions.

License and upstream attribution

This repository preserves the pinned TimeLens2 source notice in LICENSE-TIMELENS2.txt, including its academic-use and geographic restrictions. The upstream HF model card labels TimeLens2-4B as Apache-2.0; that model metadata and the pinned source notice are distinct. This release does not resolve their applicability or grant an additional unrestricted MIT/Apache license to CTVG-4B. Review the applicable upstream terms and obtain any required permissions before use or redistribution; public availability alone does not imply permission for commercial use.

We thank the authors of TimeLens2 and OMTG for releasing their code and training datasets. The work also builds on Qwen3-VL, Transformers, PEFT, PyTorch, and OMTG's vendored verl integration. Their resources retain their respective terms.

Citation

@article{chen2026grounding,
  title={Grounding with Confidence: Controllable Generative Video Temporal Grounding},
  author={Chen, Jinhao and Cui, Benlei and Jia, Ruijian and Wang, Ziheng and Wo, Tianyu and Sun, Pengfei and Huang, Longtao and Xue, Hui and Yang, Yitong and Hong, Haiwen},
  journal={arXiv preprint arXiv:2609.39883},
  year={2026}
}
Downloads last month
-
Safetensors
Model size
5B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Alibaba-VELLDEPTH/CTVG-4B

Finetuned
(1)
this model

Datasets used to train Alibaba-VELLDEPTH/CTVG-4B

Paper for Alibaba-VELLDEPTH/CTVG-4B