AI evaluation benchmark

PhysicsMind

A simulation-and-real-world benchmark for testing physical reasoning and prediction in vision-language and world models.

Year
2026
Status
active
Focus
Python · VLM Evaluation · Video Generation

What it does

PhysicsMind evaluates whether foundation models follow basic mechanics instead of relying only on visual appearance. It combines simulated and real environments and covers both visual question answering and video prediction.

The benchmark focuses on center of mass, lever equilibrium, and Newton’s first law. Its generation track checks whether predicted motion remains consistent with the physical constraints present in reference scenes.

Evidence and scope

This page is a concise guide to the work. Use the primary paper for the experimental setup, quantitative results, limitations, and formal claims; use the repository for the implementation and current reproduction instructions.