Video-Text-to-Text
Safetensors
PEFT
English
custom
lora
video-evaluation
physical-ai
world-model
judge
qwen3.5
armanakbari4 commited on
Commit
a1a8d98
·
verified ·
1 Parent(s): dd29a8a

infer.py: render the exact eval prompts, disable thinking, pass video metadata

Browse files

- Sub-questions use the q1:/q2: numbering and the per-law questions from the PhyGround repo (evals/sub_questions.py); criterion text matches evals/physics_criteria.py. Rendered prompts are now identical to .
- Qwen3.5 thinking is turned off (enable_thinking=False); the judge answers with bare JSON.
- Pass video_metadata from qwen-vl-utils>=0.0.14 so frame timestamps are correct (previously defaulted to 24 fps) and fix the crash on transformers 5.x from the per-video fps list.
- README: require qwen-vl-utils>=0.0.14 and explain how infer.py relates to the paper's vLLM pipeline.

Context: https://github.com/NU-World-Model-Embodied-AI/PhyGround/issues/1

Files changed (2) hide show
  1. README.md +11 -4
  2. infer.py +59 -52
README.md CHANGED
@@ -73,15 +73,16 @@ dimensions plus 2–4 laws associated with the prompt.
73
  pip install -U \
74
  "transformers>=5.2.0" \
75
  "peft>=0.19.1" accelerate pyyaml \
76
- "qwen-vl-utils[decord]" \
77
  huggingface_hub
78
 
79
  hf download NU-World-Model-Embodied-AI/phyjudge-9B \
80
  --local-dir ./phyjudge-9B
81
  ```
82
 
83
- The card's prompt/parser path was checked with Transformers 5.2.0 and PEFT
84
- 0.19.1. Loading the 9B base model in bf16 needs roughly 24 GB of GPU memory in the
 
85
  authors' setup. Actual memory use depends on software versions, video length,
86
  frame sampling, and resolution. Reduce `--fps` or `--max-pixels` if needed;
87
  doing so can also change the score.
@@ -110,7 +111,13 @@ python ./phyjudge-9B/infer.py \
110
 
111
  The script loads the base model recorded in `adapter_config.json`, attaches the
112
  LoRA adapter, samples the video at 2 FPS by default, and performs deterministic
113
- decoding. Output is a JSON object:
 
 
 
 
 
 
114
 
115
  ```json
116
  {
 
73
  pip install -U \
74
  "transformers>=5.2.0" \
75
  "peft>=0.19.1" accelerate pyyaml \
76
+ "qwen-vl-utils[decord]>=0.0.14" \
77
  huggingface_hub
78
 
79
  hf download NU-World-Model-Embodied-AI/phyjudge-9B \
80
  --local-dir ./phyjudge-9B
81
  ```
82
 
83
+ `infer.py` was checked end to end (video in, score out) with Transformers
84
+ 5.2.0 and qwen-vl-utils 0.0.14; the adapter was saved with PEFT 0.19.1.
85
+ Loading the 9B base model in bf16 needs roughly 24 GB of GPU memory in the
86
  authors' setup. Actual memory use depends on software versions, video length,
87
  frame sampling, and resolution. Reduce `--fps` or `--max-pixels` if needed;
88
  doing so can also change the score.
 
111
 
112
  The script loads the base model recorded in `adapter_config.json`, attaches the
113
  LoRA adapter, samples the video at 2 FPS by default, and performs deterministic
114
+ decoding with Qwen3.5 thinking turned off. It renders exactly the same prompts
115
+ as `python -m evals.vlm_eval --prompt_config subq+human.yaml` in the
116
+ [evaluation code](https://github.com/NU-World-Model-Embodied-AI/PhyGround).
117
+ The paper's numbers come from that repo's vLLM pipeline
118
+ (`scripts/score_videos.sh`), which samples at temperature 1.0, so a single
119
+ `infer.py` score can differ from it; use the GitHub pipeline to reproduce the
120
+ benchmark. Output is a JSON object:
121
 
122
  ```json
123
  {
infer.py CHANGED
@@ -32,6 +32,10 @@ from peft import PeftModel
32
  from transformers import AutoProcessor
33
 
34
 
 
 
 
 
35
  GENERAL_SUB_QUESTIONS: dict[str, list[str]] = {
36
  "SA": [
37
  "Are the main objects in the caption present in the video?",
@@ -58,12 +62,12 @@ PHYSICAL_CRITERIA: dict[str, str] = {
58
  "gravity": "Do unsupported objects fall downward? Do thrown objects follow a curved trajectory? Does poured liquid fall with gravity?",
59
  "inertia": "Do stationary objects remain still unless acted upon? Do moving objects maintain their motion unless stopped by friction, collision, or an obstacle?",
60
  "momentum": "After collision, push, or pull, is the direction of motion reasonable? Ignore speed magnitude.",
61
- "impenetrability": "Do objects maintain impenetrability -- no passing through each other?",
62
  "collision": "After impact, is there reasonable bounce/shatter/deformation? Does response match impact force?",
63
  "material": "Does each material respond according to its properties? (glass shatters, rubber bounces, metal is rigid, cloth deforms softly, etc.)",
64
  "buoyancy": "Do dense objects sink? Do wood/plastic float?",
65
  "displacement": "When you add more liquid or put an object into it, does the liquid level rise in a realistic way? Does it overflow when full?",
66
- "flow_dynamics": "Does the liquid's overall motion behave realistically over time -- flowing along surfaces, spreading, draining naturally?",
67
  "boundary_interaction": "When the liquid hits a boundary such as a rock face, container wall, or floor, does it respond realistically? Do local splash, rebound, or split patterns on impact look physically plausible?",
68
  "fluid_continuity": "Does the liquid avoid disappearing or appearing out of nowhere? Small splashes that briefly break apart are okay.",
69
  "reflection": "Does the reflection roughly match objects and colors in the scene, and avoid completely unrelated content?",
@@ -73,69 +77,62 @@ PHYSICAL_CRITERIA: dict[str, str] = {
73
 
74
  PHYSICAL_SUB_QUESTIONS: dict[str, list[str]] = {
75
  "gravity": [
76
- "Do unsupported objects or liquids move downward over time?",
77
- "Do thrown or falling objects follow a plausible gravity-driven path?",
78
- "Does the video avoid objects floating or rising without support?",
79
  ],
80
  "inertia": [
81
- "Do stationary objects remain still unless a visible force acts on them?",
82
- "Do moving objects continue plausibly until friction, collision, or an obstacle changes their motion?",
83
- "Does the video avoid unexplained starts, stops, or direction changes?",
84
  ],
85
  "momentum": [
86
- "After contact, push, pull, or collision, are motion directions plausible?",
87
- "Does the reacting object move in a direction consistent with the interaction?",
88
- "Does the video avoid impossible reversals or unrelated motion changes?",
89
  ],
90
  "impenetrability": [
91
- "Do solid objects avoid passing through one another?",
92
- "Do contacts and overlaps remain physically plausible?",
93
- "Does the video avoid obvious clipping or penetration artifacts?",
94
  ],
95
  "collision": [
96
- "Does impact cause a plausible bounce, break, deformation, or transfer of motion?",
97
- "Is the response direction consistent with the collision?",
98
- "Does the response avoid being much too weak, too strong, or unrelated to the impact?",
99
  ],
100
  "material": [
101
- "Do objects respond consistently with their apparent material?",
102
- "Are rigid, soft, brittle, elastic, or fluid-like objects animated appropriately?",
103
- "Does the video avoid material behavior that contradicts the scene?",
104
  ],
105
  "buoyancy": [
106
- "Do objects sink or float in a way consistent with apparent density?",
107
- "Does the floating or sinking behavior stay stable over time?",
108
- "Does the video avoid unsupported hovering or impossible underwater motion?",
109
  ],
110
  "displacement": [
111
- "Does liquid level rise when volume is added or an object enters it?",
112
- "Does overflow happen only when the container is plausibly full?",
113
- "Does the liquid volume remain visually plausible?",
114
  ],
115
  "flow_dynamics": [
116
- "Does liquid flow along surfaces, spread, or drain naturally?",
117
- "Does the flow direction follow gravity and boundaries?",
118
- "Does the video avoid abrupt stops, reversals, or unsupported uphill flow?",
119
  ],
120
  "boundary_interaction": [
121
- "Does liquid react plausibly when hitting a wall, floor, container, or obstacle?",
122
- "Are splash, rebound, or split patterns locally plausible?",
123
- "Does the liquid remain consistent after interacting with boundaries?",
124
  ],
125
  "fluid_continuity": [
126
- "Does liquid avoid disappearing or appearing without cause?",
127
- "Does the amount of liquid remain broadly consistent?",
128
- "Are splashes and separations temporary and physically plausible?",
129
  ],
130
  "reflection": [
131
- "Does the reflection match nearby objects, colors, and motion?",
132
- "Does the reflected content stay spatially consistent with the scene?",
133
- "Does the video avoid unrelated or impossible reflection content?",
134
  ],
135
  "shadow": [
136
- "Are shadows consistent with the apparent light source direction?",
137
- "Do shadows move with the objects that cast them?",
138
- "Does the video avoid missing, detached, or contradictory shadows?",
139
  ],
140
  }
141
 
@@ -151,7 +148,7 @@ def load_yaml(path: Path) -> dict[str, Any]:
151
 
152
 
153
  def questions_block(questions: list[str]) -> str:
154
- return "\n".join(f"{idx}. {question}" for idx, question in enumerate(questions, 1))
155
 
156
 
157
  def build_prompt(
@@ -257,18 +254,21 @@ def prepare_inputs(
257
  fps: float,
258
  max_pixels: int,
259
  ) -> dict[str, Any]:
 
 
260
  text = processor.apply_chat_template(
261
  messages,
262
  tokenize=False,
263
  add_generation_prompt=True,
 
264
  )
265
 
266
  try:
267
  from qwen_vl_utils import process_vision_info
268
  except ImportError as exc:
269
  raise ImportError(
270
- "qwen-vl-utils is required for local video inference. "
271
- "Install it with: pip install qwen-vl-utils[decord]"
272
  ) from exc
273
 
274
  for msg in messages:
@@ -279,19 +279,26 @@ def prepare_inputs(
279
  item.setdefault("fps", fps)
280
  item.setdefault("max_pixels", max_pixels)
281
 
282
- try:
283
- image_inputs, video_inputs, video_kwargs = process_vision_info(
284
- messages,
285
- return_video_kwargs=True,
286
- )
287
- except TypeError:
288
- image_inputs, video_inputs = process_vision_info(messages)
289
- video_kwargs = {}
 
 
 
 
 
290
 
291
  inputs = processor(
292
  text=[text],
293
  images=image_inputs,
294
  videos=video_inputs,
 
 
295
  padding=True,
296
  return_tensors="pt",
297
  **video_kwargs,
 
32
  from transformers import AutoProcessor
33
 
34
 
35
+ # The tables below mirror evals/prompts/__init__.py (GENERAL_SUB_QUESTIONS),
36
+ # evals/physics_criteria.py (CRITERIA) and evals/sub_questions.py (SUB_QUESTIONS)
37
+ # in https://github.com/NU-World-Model-Embodied-AI/PhyGround, so the rendered
38
+ # prompts match `python -m evals.vlm_eval --prompt_config subq+human.yaml`.
39
  GENERAL_SUB_QUESTIONS: dict[str, list[str]] = {
40
  "SA": [
41
  "Are the main objects in the caption present in the video?",
 
62
  "gravity": "Do unsupported objects fall downward? Do thrown objects follow a curved trajectory? Does poured liquid fall with gravity?",
63
  "inertia": "Do stationary objects remain still unless acted upon? Do moving objects maintain their motion unless stopped by friction, collision, or an obstacle?",
64
  "momentum": "After collision, push, or pull, is the direction of motion reasonable? Ignore speed magnitude.",
65
+ "impenetrability": "Do objects maintain impenetrability — no passing through each other?",
66
  "collision": "After impact, is there reasonable bounce/shatter/deformation? Does response match impact force?",
67
  "material": "Does each material respond according to its properties? (glass shatters, rubber bounces, metal is rigid, cloth deforms softly, etc.)",
68
  "buoyancy": "Do dense objects sink? Do wood/plastic float?",
69
  "displacement": "When you add more liquid or put an object into it, does the liquid level rise in a realistic way? Does it overflow when full?",
70
+ "flow_dynamics": "Does the liquid's overall motion behave realistically over time — flowing along surfaces, spreading, draining naturally?",
71
  "boundary_interaction": "When the liquid hits a boundary such as a rock face, container wall, or floor, does it respond realistically? Do local splash, rebound, or split patterns on impact look physically plausible?",
72
  "fluid_continuity": "Does the liquid avoid disappearing or appearing out of nowhere? Small splashes that briefly break apart are okay.",
73
  "reflection": "Does the reflection roughly match objects and colors in the scene, and avoid completely unrelated content?",
 
77
 
78
  PHYSICAL_SUB_QUESTIONS: dict[str, list[str]] = {
79
  "gravity": [
80
+ "Do any unsupported objects float upward or hover in mid-air?",
81
+ "Does any heavy object drift or float down unrealistically slowly, as if it had no weight?",
82
+ "Does any object fall sideways or in an unnatural direction instead of downward?",
83
  ],
84
  "inertia": [
85
+ "Does any object spontaneously start moving without a visible cause?",
86
+ "Does any moving object suddenly stop or reverse direction with no visible contact or force?",
 
87
  ],
88
  "momentum": [
89
+ "Does any object fly off in a direction completely unrelated to how it was hit or pushed?",
90
+ "Does the recoil direction contradict the direction of the applied force?",
 
91
  ],
92
  "impenetrability": [
93
+ "Does any object pass through another solid object as if it were not there?",
94
+ "Does any part of an object clip into a wall, floor, or another object's interior?",
 
95
  ],
96
  "collision": [
97
+ "Does any object remain completely unaffected after being clearly hit by another object?",
98
+ "Does the collision response look wildly too weak or too strong for the visible impact?",
99
+ "Does any object shatter or deform dramatically from a very light touch with no reasonable force?",
100
  ],
101
  "material": [
102
+ "Does any rigid material (glass, metal, stone) bend or stretch like rubber?",
103
+ "Does any soft material (cloth, rubber, rope) behave as if it were completely rigid and unbending?",
 
104
  ],
105
  "buoyancy": [
106
+ "Does any heavy object (metal, stone) float on the liquid surface?",
107
+ "Does any light object (wood, cork, plastic) sink to the bottom?",
 
108
  ],
109
  "displacement": [
110
+ "Does the liquid level remain completely unchanged when a large object is submerged?",
111
+ "Does a full container spill or overflow when more liquid or an object is added?",
112
+ "Does the liquid level behave in a clearly impossible way, such as dropping when volume is added?",
113
  ],
114
  "flow_dynamics": [
115
+ "Does liquid flow uphill or against gravity without any force pushing it?",
116
+ "On a flat surface, does liquid spread outward and become thinner over time?",
117
+ "Does liquid suddenly stop flowing or freeze in place without an obvious reason?",
118
  ],
119
  "boundary_interaction": [
120
+ "Does liquid ignore a solid boundary and continue moving as if nothing were there?",
121
+ "Does liquid striking a surface produce visible droplets, spray, or ripples?",
122
+ "Does liquid accumulate or pool on the wrong side of a barrier?",
123
  ],
124
  "fluid_continuity": [
125
+ "Does a continuous pour or flow stay connected as a stream without sudden gaps?",
126
+ "Does liquid disappear into nothing (not counting brief splashes)?",
127
+ "Does liquid appear from nowhere with no visible source?",
128
  ],
129
  "reflection": [
130
+ "Do reflections show completely unrelated content not present in the scene?",
131
+ "Does a reflection remain completely static while the scene clearly changes around it?",
 
132
  ],
133
  "shadow": [
134
+ "Do different shadows in the same scene point in contradictory directions?",
135
+ "Does any shadow remain fixed in place while its object clearly moves?",
 
136
  ],
137
  }
138
 
 
148
 
149
 
150
  def questions_block(questions: list[str]) -> str:
151
+ return "\n".join(f"q{idx}: {question}" for idx, question in enumerate(questions, 1))
152
 
153
 
154
  def build_prompt(
 
254
  fps: float,
255
  max_pixels: int,
256
  ) -> dict[str, Any]:
257
+ # Qwen3.5 thinks by default; the judge answers with bare JSON, so turn it off
258
+ # (same as `evals.vlm_eval --no-thinking`).
259
  text = processor.apply_chat_template(
260
  messages,
261
  tokenize=False,
262
  add_generation_prompt=True,
263
+ enable_thinking=False,
264
  )
265
 
266
  try:
267
  from qwen_vl_utils import process_vision_info
268
  except ImportError as exc:
269
  raise ImportError(
270
+ "qwen-vl-utils>=0.0.14 is required for local video inference. "
271
+ "Install it with: pip install 'qwen-vl-utils[decord]>=0.0.14'"
272
  ) from exc
273
 
274
  for msg in messages:
 
279
  item.setdefault("fps", fps)
280
  item.setdefault("max_pixels", max_pixels)
281
 
282
+ # Qwen3.5 uses the Qwen3-VL video processor: frames are resized to multiples
283
+ # of its patch size (16), and the processor needs per-video metadata to put
284
+ # the right frame timestamps in the prompt (without it, it assumes 24 fps).
285
+ patch_size = getattr(getattr(processor, "video_processor", None), "patch_size", 16)
286
+ image_inputs, video_inputs, video_kwargs = process_vision_info(
287
+ messages,
288
+ return_video_kwargs=True,
289
+ return_video_metadata=True,
290
+ image_patch_size=patch_size,
291
+ )
292
+ video_metadata = None
293
+ if video_inputs is not None:
294
+ video_inputs, video_metadata = (list(x) for x in zip(*video_inputs))
295
 
296
  inputs = processor(
297
  text=[text],
298
  images=image_inputs,
299
  videos=video_inputs,
300
+ video_metadata=video_metadata,
301
+ do_resize=False,
302
  padding=True,
303
  return_tensors="pt",
304
  **video_kwargs,