Video-Text-to-Text
Safetensors
PEFT
English
custom
lora
video-evaluation
physical-ai
world-model
judge
qwen3.5
Instructions to use NU-World-Model-Embodied-AI/phyjudge-9B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use NU-World-Model-Embodied-AI/phyjudge-9B with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
infer.py: render the exact eval prompts, disable thinking, pass video metadata
Browse files- Sub-questions use the q1:/q2: numbering and the per-law questions from the PhyGround repo (evals/sub_questions.py); criterion text matches evals/physics_criteria.py. Rendered prompts are now identical to .
- Qwen3.5 thinking is turned off (enable_thinking=False); the judge answers with bare JSON.
- Pass video_metadata from qwen-vl-utils>=0.0.14 so frame timestamps are correct (previously defaulted to 24 fps) and fix the crash on transformers 5.x from the per-video fps list.
- README: require qwen-vl-utils>=0.0.14 and explain how infer.py relates to the paper's vLLM pipeline.
Context: https://github.com/NU-World-Model-Embodied-AI/PhyGround/issues/1
README.md
CHANGED
|
@@ -73,15 +73,16 @@ dimensions plus 2–4 laws associated with the prompt.
|
|
| 73 |
pip install -U \
|
| 74 |
"transformers>=5.2.0" \
|
| 75 |
"peft>=0.19.1" accelerate pyyaml \
|
| 76 |
-
"qwen-vl-utils[decord]" \
|
| 77 |
huggingface_hub
|
| 78 |
|
| 79 |
hf download NU-World-Model-Embodied-AI/phyjudge-9B \
|
| 80 |
--local-dir ./phyjudge-9B
|
| 81 |
```
|
| 82 |
|
| 83 |
-
|
| 84 |
-
|
|
|
|
| 85 |
authors' setup. Actual memory use depends on software versions, video length,
|
| 86 |
frame sampling, and resolution. Reduce `--fps` or `--max-pixels` if needed;
|
| 87 |
doing so can also change the score.
|
|
@@ -110,7 +111,13 @@ python ./phyjudge-9B/infer.py \
|
|
| 110 |
|
| 111 |
The script loads the base model recorded in `adapter_config.json`, attaches the
|
| 112 |
LoRA adapter, samples the video at 2 FPS by default, and performs deterministic
|
| 113 |
-
decoding.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 114 |
|
| 115 |
```json
|
| 116 |
{
|
|
|
|
| 73 |
pip install -U \
|
| 74 |
"transformers>=5.2.0" \
|
| 75 |
"peft>=0.19.1" accelerate pyyaml \
|
| 76 |
+
"qwen-vl-utils[decord]>=0.0.14" \
|
| 77 |
huggingface_hub
|
| 78 |
|
| 79 |
hf download NU-World-Model-Embodied-AI/phyjudge-9B \
|
| 80 |
--local-dir ./phyjudge-9B
|
| 81 |
```
|
| 82 |
|
| 83 |
+
`infer.py` was checked end to end (video in, score out) with Transformers
|
| 84 |
+
5.2.0 and qwen-vl-utils 0.0.14; the adapter was saved with PEFT 0.19.1.
|
| 85 |
+
Loading the 9B base model in bf16 needs roughly 24 GB of GPU memory in the
|
| 86 |
authors' setup. Actual memory use depends on software versions, video length,
|
| 87 |
frame sampling, and resolution. Reduce `--fps` or `--max-pixels` if needed;
|
| 88 |
doing so can also change the score.
|
|
|
|
| 111 |
|
| 112 |
The script loads the base model recorded in `adapter_config.json`, attaches the
|
| 113 |
LoRA adapter, samples the video at 2 FPS by default, and performs deterministic
|
| 114 |
+
decoding with Qwen3.5 thinking turned off. It renders exactly the same prompts
|
| 115 |
+
as `python -m evals.vlm_eval --prompt_config subq+human.yaml` in the
|
| 116 |
+
[evaluation code](https://github.com/NU-World-Model-Embodied-AI/PhyGround).
|
| 117 |
+
The paper's numbers come from that repo's vLLM pipeline
|
| 118 |
+
(`scripts/score_videos.sh`), which samples at temperature 1.0, so a single
|
| 119 |
+
`infer.py` score can differ from it; use the GitHub pipeline to reproduce the
|
| 120 |
+
benchmark. Output is a JSON object:
|
| 121 |
|
| 122 |
```json
|
| 123 |
{
|
infer.py
CHANGED
|
@@ -32,6 +32,10 @@ from peft import PeftModel
|
|
| 32 |
from transformers import AutoProcessor
|
| 33 |
|
| 34 |
|
|
|
|
|
|
|
|
|
|
|
|
|
| 35 |
GENERAL_SUB_QUESTIONS: dict[str, list[str]] = {
|
| 36 |
"SA": [
|
| 37 |
"Are the main objects in the caption present in the video?",
|
|
@@ -58,12 +62,12 @@ PHYSICAL_CRITERIA: dict[str, str] = {
|
|
| 58 |
"gravity": "Do unsupported objects fall downward? Do thrown objects follow a curved trajectory? Does poured liquid fall with gravity?",
|
| 59 |
"inertia": "Do stationary objects remain still unless acted upon? Do moving objects maintain their motion unless stopped by friction, collision, or an obstacle?",
|
| 60 |
"momentum": "After collision, push, or pull, is the direction of motion reasonable? Ignore speed magnitude.",
|
| 61 |
-
"impenetrability": "Do objects maintain impenetrability
|
| 62 |
"collision": "After impact, is there reasonable bounce/shatter/deformation? Does response match impact force?",
|
| 63 |
"material": "Does each material respond according to its properties? (glass shatters, rubber bounces, metal is rigid, cloth deforms softly, etc.)",
|
| 64 |
"buoyancy": "Do dense objects sink? Do wood/plastic float?",
|
| 65 |
"displacement": "When you add more liquid or put an object into it, does the liquid level rise in a realistic way? Does it overflow when full?",
|
| 66 |
-
"flow_dynamics": "Does the liquid's overall motion behave realistically over time
|
| 67 |
"boundary_interaction": "When the liquid hits a boundary such as a rock face, container wall, or floor, does it respond realistically? Do local splash, rebound, or split patterns on impact look physically plausible?",
|
| 68 |
"fluid_continuity": "Does the liquid avoid disappearing or appearing out of nowhere? Small splashes that briefly break apart are okay.",
|
| 69 |
"reflection": "Does the reflection roughly match objects and colors in the scene, and avoid completely unrelated content?",
|
|
@@ -73,69 +77,62 @@ PHYSICAL_CRITERIA: dict[str, str] = {
|
|
| 73 |
|
| 74 |
PHYSICAL_SUB_QUESTIONS: dict[str, list[str]] = {
|
| 75 |
"gravity": [
|
| 76 |
-
"Do unsupported objects
|
| 77 |
-
"
|
| 78 |
-
"Does
|
| 79 |
],
|
| 80 |
"inertia": [
|
| 81 |
-
"
|
| 82 |
-
"
|
| 83 |
-
"Does the video avoid unexplained starts, stops, or direction changes?",
|
| 84 |
],
|
| 85 |
"momentum": [
|
| 86 |
-
"
|
| 87 |
-
"Does the
|
| 88 |
-
"Does the video avoid impossible reversals or unrelated motion changes?",
|
| 89 |
],
|
| 90 |
"impenetrability": [
|
| 91 |
-
"
|
| 92 |
-
"
|
| 93 |
-
"Does the video avoid obvious clipping or penetration artifacts?",
|
| 94 |
],
|
| 95 |
"collision": [
|
| 96 |
-
"Does
|
| 97 |
-
"
|
| 98 |
-
"Does
|
| 99 |
],
|
| 100 |
"material": [
|
| 101 |
-
"
|
| 102 |
-
"
|
| 103 |
-
"Does the video avoid material behavior that contradicts the scene?",
|
| 104 |
],
|
| 105 |
"buoyancy": [
|
| 106 |
-
"
|
| 107 |
-
"Does
|
| 108 |
-
"Does the video avoid unsupported hovering or impossible underwater motion?",
|
| 109 |
],
|
| 110 |
"displacement": [
|
| 111 |
-
"Does liquid level
|
| 112 |
-
"Does
|
| 113 |
-
"Does the liquid
|
| 114 |
],
|
| 115 |
"flow_dynamics": [
|
| 116 |
-
"Does liquid flow
|
| 117 |
-
"
|
| 118 |
-
"Does
|
| 119 |
],
|
| 120 |
"boundary_interaction": [
|
| 121 |
-
"Does liquid
|
| 122 |
-
"
|
| 123 |
-
"Does
|
| 124 |
],
|
| 125 |
"fluid_continuity": [
|
| 126 |
-
"Does
|
| 127 |
-
"Does
|
| 128 |
-
"
|
| 129 |
],
|
| 130 |
"reflection": [
|
| 131 |
-
"
|
| 132 |
-
"Does
|
| 133 |
-
"Does the video avoid unrelated or impossible reflection content?",
|
| 134 |
],
|
| 135 |
"shadow": [
|
| 136 |
-
"
|
| 137 |
-
"
|
| 138 |
-
"Does the video avoid missing, detached, or contradictory shadows?",
|
| 139 |
],
|
| 140 |
}
|
| 141 |
|
|
@@ -151,7 +148,7 @@ def load_yaml(path: Path) -> dict[str, Any]:
|
|
| 151 |
|
| 152 |
|
| 153 |
def questions_block(questions: list[str]) -> str:
|
| 154 |
-
return "\n".join(f"{idx}
|
| 155 |
|
| 156 |
|
| 157 |
def build_prompt(
|
|
@@ -257,18 +254,21 @@ def prepare_inputs(
|
|
| 257 |
fps: float,
|
| 258 |
max_pixels: int,
|
| 259 |
) -> dict[str, Any]:
|
|
|
|
|
|
|
| 260 |
text = processor.apply_chat_template(
|
| 261 |
messages,
|
| 262 |
tokenize=False,
|
| 263 |
add_generation_prompt=True,
|
|
|
|
| 264 |
)
|
| 265 |
|
| 266 |
try:
|
| 267 |
from qwen_vl_utils import process_vision_info
|
| 268 |
except ImportError as exc:
|
| 269 |
raise ImportError(
|
| 270 |
-
"qwen-vl-utils is required for local video inference. "
|
| 271 |
-
"Install it with: pip install qwen-vl-utils[decord]"
|
| 272 |
) from exc
|
| 273 |
|
| 274 |
for msg in messages:
|
|
@@ -279,19 +279,26 @@ def prepare_inputs(
|
|
| 279 |
item.setdefault("fps", fps)
|
| 280 |
item.setdefault("max_pixels", max_pixels)
|
| 281 |
|
| 282 |
-
|
| 283 |
-
|
| 284 |
-
|
| 285 |
-
|
| 286 |
-
|
| 287 |
-
|
| 288 |
-
|
| 289 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 290 |
|
| 291 |
inputs = processor(
|
| 292 |
text=[text],
|
| 293 |
images=image_inputs,
|
| 294 |
videos=video_inputs,
|
|
|
|
|
|
|
| 295 |
padding=True,
|
| 296 |
return_tensors="pt",
|
| 297 |
**video_kwargs,
|
|
|
|
| 32 |
from transformers import AutoProcessor
|
| 33 |
|
| 34 |
|
| 35 |
+
# The tables below mirror evals/prompts/__init__.py (GENERAL_SUB_QUESTIONS),
|
| 36 |
+
# evals/physics_criteria.py (CRITERIA) and evals/sub_questions.py (SUB_QUESTIONS)
|
| 37 |
+
# in https://github.com/NU-World-Model-Embodied-AI/PhyGround, so the rendered
|
| 38 |
+
# prompts match `python -m evals.vlm_eval --prompt_config subq+human.yaml`.
|
| 39 |
GENERAL_SUB_QUESTIONS: dict[str, list[str]] = {
|
| 40 |
"SA": [
|
| 41 |
"Are the main objects in the caption present in the video?",
|
|
|
|
| 62 |
"gravity": "Do unsupported objects fall downward? Do thrown objects follow a curved trajectory? Does poured liquid fall with gravity?",
|
| 63 |
"inertia": "Do stationary objects remain still unless acted upon? Do moving objects maintain their motion unless stopped by friction, collision, or an obstacle?",
|
| 64 |
"momentum": "After collision, push, or pull, is the direction of motion reasonable? Ignore speed magnitude.",
|
| 65 |
+
"impenetrability": "Do objects maintain impenetrability — no passing through each other?",
|
| 66 |
"collision": "After impact, is there reasonable bounce/shatter/deformation? Does response match impact force?",
|
| 67 |
"material": "Does each material respond according to its properties? (glass shatters, rubber bounces, metal is rigid, cloth deforms softly, etc.)",
|
| 68 |
"buoyancy": "Do dense objects sink? Do wood/plastic float?",
|
| 69 |
"displacement": "When you add more liquid or put an object into it, does the liquid level rise in a realistic way? Does it overflow when full?",
|
| 70 |
+
"flow_dynamics": "Does the liquid's overall motion behave realistically over time — flowing along surfaces, spreading, draining naturally?",
|
| 71 |
"boundary_interaction": "When the liquid hits a boundary such as a rock face, container wall, or floor, does it respond realistically? Do local splash, rebound, or split patterns on impact look physically plausible?",
|
| 72 |
"fluid_continuity": "Does the liquid avoid disappearing or appearing out of nowhere? Small splashes that briefly break apart are okay.",
|
| 73 |
"reflection": "Does the reflection roughly match objects and colors in the scene, and avoid completely unrelated content?",
|
|
|
|
| 77 |
|
| 78 |
PHYSICAL_SUB_QUESTIONS: dict[str, list[str]] = {
|
| 79 |
"gravity": [
|
| 80 |
+
"Do any unsupported objects float upward or hover in mid-air?",
|
| 81 |
+
"Does any heavy object drift or float down unrealistically slowly, as if it had no weight?",
|
| 82 |
+
"Does any object fall sideways or in an unnatural direction instead of downward?",
|
| 83 |
],
|
| 84 |
"inertia": [
|
| 85 |
+
"Does any object spontaneously start moving without a visible cause?",
|
| 86 |
+
"Does any moving object suddenly stop or reverse direction with no visible contact or force?",
|
|
|
|
| 87 |
],
|
| 88 |
"momentum": [
|
| 89 |
+
"Does any object fly off in a direction completely unrelated to how it was hit or pushed?",
|
| 90 |
+
"Does the recoil direction contradict the direction of the applied force?",
|
|
|
|
| 91 |
],
|
| 92 |
"impenetrability": [
|
| 93 |
+
"Does any object pass through another solid object as if it were not there?",
|
| 94 |
+
"Does any part of an object clip into a wall, floor, or another object's interior?",
|
|
|
|
| 95 |
],
|
| 96 |
"collision": [
|
| 97 |
+
"Does any object remain completely unaffected after being clearly hit by another object?",
|
| 98 |
+
"Does the collision response look wildly too weak or too strong for the visible impact?",
|
| 99 |
+
"Does any object shatter or deform dramatically from a very light touch with no reasonable force?",
|
| 100 |
],
|
| 101 |
"material": [
|
| 102 |
+
"Does any rigid material (glass, metal, stone) bend or stretch like rubber?",
|
| 103 |
+
"Does any soft material (cloth, rubber, rope) behave as if it were completely rigid and unbending?",
|
|
|
|
| 104 |
],
|
| 105 |
"buoyancy": [
|
| 106 |
+
"Does any heavy object (metal, stone) float on the liquid surface?",
|
| 107 |
+
"Does any light object (wood, cork, plastic) sink to the bottom?",
|
|
|
|
| 108 |
],
|
| 109 |
"displacement": [
|
| 110 |
+
"Does the liquid level remain completely unchanged when a large object is submerged?",
|
| 111 |
+
"Does a full container spill or overflow when more liquid or an object is added?",
|
| 112 |
+
"Does the liquid level behave in a clearly impossible way, such as dropping when volume is added?",
|
| 113 |
],
|
| 114 |
"flow_dynamics": [
|
| 115 |
+
"Does liquid flow uphill or against gravity without any force pushing it?",
|
| 116 |
+
"On a flat surface, does liquid spread outward and become thinner over time?",
|
| 117 |
+
"Does liquid suddenly stop flowing or freeze in place without an obvious reason?",
|
| 118 |
],
|
| 119 |
"boundary_interaction": [
|
| 120 |
+
"Does liquid ignore a solid boundary and continue moving as if nothing were there?",
|
| 121 |
+
"Does liquid striking a surface produce visible droplets, spray, or ripples?",
|
| 122 |
+
"Does liquid accumulate or pool on the wrong side of a barrier?",
|
| 123 |
],
|
| 124 |
"fluid_continuity": [
|
| 125 |
+
"Does a continuous pour or flow stay connected as a stream without sudden gaps?",
|
| 126 |
+
"Does liquid disappear into nothing (not counting brief splashes)?",
|
| 127 |
+
"Does liquid appear from nowhere with no visible source?",
|
| 128 |
],
|
| 129 |
"reflection": [
|
| 130 |
+
"Do reflections show completely unrelated content not present in the scene?",
|
| 131 |
+
"Does a reflection remain completely static while the scene clearly changes around it?",
|
|
|
|
| 132 |
],
|
| 133 |
"shadow": [
|
| 134 |
+
"Do different shadows in the same scene point in contradictory directions?",
|
| 135 |
+
"Does any shadow remain fixed in place while its object clearly moves?",
|
|
|
|
| 136 |
],
|
| 137 |
}
|
| 138 |
|
|
|
|
| 148 |
|
| 149 |
|
| 150 |
def questions_block(questions: list[str]) -> str:
|
| 151 |
+
return "\n".join(f"q{idx}: {question}" for idx, question in enumerate(questions, 1))
|
| 152 |
|
| 153 |
|
| 154 |
def build_prompt(
|
|
|
|
| 254 |
fps: float,
|
| 255 |
max_pixels: int,
|
| 256 |
) -> dict[str, Any]:
|
| 257 |
+
# Qwen3.5 thinks by default; the judge answers with bare JSON, so turn it off
|
| 258 |
+
# (same as `evals.vlm_eval --no-thinking`).
|
| 259 |
text = processor.apply_chat_template(
|
| 260 |
messages,
|
| 261 |
tokenize=False,
|
| 262 |
add_generation_prompt=True,
|
| 263 |
+
enable_thinking=False,
|
| 264 |
)
|
| 265 |
|
| 266 |
try:
|
| 267 |
from qwen_vl_utils import process_vision_info
|
| 268 |
except ImportError as exc:
|
| 269 |
raise ImportError(
|
| 270 |
+
"qwen-vl-utils>=0.0.14 is required for local video inference. "
|
| 271 |
+
"Install it with: pip install 'qwen-vl-utils[decord]>=0.0.14'"
|
| 272 |
) from exc
|
| 273 |
|
| 274 |
for msg in messages:
|
|
|
|
| 279 |
item.setdefault("fps", fps)
|
| 280 |
item.setdefault("max_pixels", max_pixels)
|
| 281 |
|
| 282 |
+
# Qwen3.5 uses the Qwen3-VL video processor: frames are resized to multiples
|
| 283 |
+
# of its patch size (16), and the processor needs per-video metadata to put
|
| 284 |
+
# the right frame timestamps in the prompt (without it, it assumes 24 fps).
|
| 285 |
+
patch_size = getattr(getattr(processor, "video_processor", None), "patch_size", 16)
|
| 286 |
+
image_inputs, video_inputs, video_kwargs = process_vision_info(
|
| 287 |
+
messages,
|
| 288 |
+
return_video_kwargs=True,
|
| 289 |
+
return_video_metadata=True,
|
| 290 |
+
image_patch_size=patch_size,
|
| 291 |
+
)
|
| 292 |
+
video_metadata = None
|
| 293 |
+
if video_inputs is not None:
|
| 294 |
+
video_inputs, video_metadata = (list(x) for x in zip(*video_inputs))
|
| 295 |
|
| 296 |
inputs = processor(
|
| 297 |
text=[text],
|
| 298 |
images=image_inputs,
|
| 299 |
videos=video_inputs,
|
| 300 |
+
video_metadata=video_metadata,
|
| 301 |
+
do_resize=False,
|
| 302 |
padding=True,
|
| 303 |
return_tensors="pt",
|
| 304 |
**video_kwargs,
|