Train a World Model in CoppeliaSim
Train a model that predicts puck motion from actions in CoppeliaSim.
Introduction
30 Second Summary
A robot can look capable until it meets a movement it never practised. That gap can turn a confident prediction into the wrong next move.
In this project, you will build an action-conditioned world model around a force-controlled puck in CoppeliaSim. A deliberate gap in its training data lets you see how balanced experience improves predictions on unseen directions.
What You'll Build
Your final browser report overlays the puck's real CoppeliaSim trajectory with the path imagined by your trained model to reveal how prediction error grows across the rollout.
By the end of this project, you'll have:
- A synchronized puck experiment that turns visible force responses into reusable datasets of state-action-next-state transitions.
- Biased and balanced PyTorch models that show how action coverage changes accuracy on unseen directions.
- A recursive rollout report that overlays real and imagined paths in the browser with multi-step RMSE.
- Secret Mission: A sim-to-sim transfer test that keeps the learned model fixed while changing the physics engine from MuJoCo to Bullet Physics.
Are there any prerequisites?
You need existing CoppeliaSim access on a Windows computer. Basic Python syntax is enough because the guide covers Python installation plus every required package.
Before We Start
Before any setup begins, lock in the experiment you are about to build. Your model observes the puck state [x, y, vx, vy]. It receives the force action [fx, fy]. It predicts the next-state delta [Δx, Δy, Δvx, Δvy].
Connect Python to CoppeliaSim
Reliable world-model data depends on repeatable transitions. Each force action needs to advance the simulator by exactly one controlled step.
You will connect Python to CoppeliaSim through the ZeroMQ Remote API. PowerShell will run the client inside an isolated environment.
In this step, get ready to:
- Install the required Python version inside an isolated environment.
- Install the pinned packages and select MuJoCo in CoppeliaSim.
- Connect to CoppeliaSim and advance one synchronized simulation step.
Check Python and isolate the environment
A virtual environment keeps this project's packages inside its own folder. This prevents another Python project from changing the versions used here.
- Open File Explorer by pressing Win+E.
- Navigate to the local project folder you chose earlier.
You should see the contents of your chosen project folder. The next PowerShell window needs to start from this location.
- Select the address bar at the top of File Explorer.
- Type PowerShell and press Enter.
- Check the available Python version by running this command:
python --version
What does this command do?
The command asks the active Python installation to print its version. This project uses Python 3.13.15.
✔️ I see the required Python version
Python 3.13.15 is ready. You can create the project's isolated environment.
ⓧ I see an older version
The existing Python installation is older than the version used by this project. Install Python 3.13.15 before creating the environment.
- Open the official Python on Windows guide.
- Install Python 3.13.15 through the Python Install Manager.
- Close the current PowerShell window after the installation finishes.
- Open a fresh PowerShell window from your project folder using the File Explorer address bar.
- Check the version again by running:
python --version
What should I see?
PowerShell should report Python 3.13.15. The fresh window ensures the new installation is available to your terminal.
ⓧ Command not found
PowerShell cannot currently find a Python installation. The Python Install Manager provides the required version on Windows.
- Open the official Python on Windows guide.
- Install Python 3.13.15 through the Python Install Manager.
- Close the current PowerShell window after the installation finishes.
- Open a fresh PowerShell window from your project folder using the File Explorer address bar.
- Confirm Python is available by running:
python --version
What should I see?
PowerShell should now report Python 3.13.15. This confirms the installation is available from your terminal.
- Create .venv inside your project folder and activate it by running these commands:
python -m venv .venv
.\.venv\Scripts\Activate.ps1
What do these commands do?
The first command creates an isolated Python environment in .venv. The second command activates that environment in the current PowerShell session.
Future package installations now belong to this project. Your other Python environments remain unchanged.
Your PowerShell prompt should now show the active environment name. This confirms that later commands use the Python installation inside .venv.
- If PowerShell blocks the activation script, apply the documented current-user policy and activate the environment by running:
Set-ExecutionPolicy -ExecutionPolicy RemoteSigned -Scope CurrentUser
.\.venv\Scripts\Activate.ps1
Why is this policy change needed?
The first command allows locally created activation scripts for your Windows user account. The second command retries the environment activation.
The current-user scope avoids changing the policy for other Windows accounts. You only need this policy command when PowerShell blocks the activation script.
Environment still inactive?
- Confirm PowerShell is running from the project folder that contains .venv.
- Confirm the environment creation command completed before you ran the activation script.
- Still stuck? Help me activate my Windows Python environment.
Install packages and prepare CoppeliaSim
The remote API package carries commands between Python and CoppeliaSim. PyTorch provides the tensors and neural-network tools used later in the project.
- Open Windows search by pressing the Windows key.
- Type Notepad and press Enter.
- Add the pinned project dependencies by pasting this content into Notepad:
coppeliasim-zmqremoteapi-client==2.0.4
torch==2.14.1
What does this file record?
The first line pins the CoppeliaSim ZeroMQ Remote API client to 2.0.4. The second line pins PyTorch to 2.14.1.
Exact versions make the environment reproducible. Anyone rebuilding the project can use the same dependency set.
You should now see both dependency lines in Notepad. The order should match the reference above.
- Open the File menu in Notepad.
- Select Save As.
- Choose the local project folder from earlier.
- Enter requirements.txt in the File name box.
- Confirm the save.
The Notepad tab should now show requirements.txt. This confirms the dependency record exists inside your project folder.
- Return to the active PowerShell window from your project folder.
- Install the two pinned packages by running these commands:
python -m pip install coppeliasim-zmqremoteapi-client==2.0.4
python -m pip install torch==2.14.1
What do these commands install?
The first command installs the CoppeliaSim ZeroMQ Remote API client 2.0.4. This client lets the Python process call the simulator's remote API.
The second command installs PyTorch 2.14.1 inside the active environment. The PyTorch package is the larger installation, so the terminal may remain busy while it downloads.
PowerShell returns to the active environment prompt after both installations finish. The completed output confirms the packages were installed inside .venv.
Package installation failed?
- Confirm the PowerShell prompt still shows the active environment name.
- Confirm the version check reports Python 3.13.15.
- Still stuck? Help me troubleshoot the pinned package installation.
CoppeliaSim must be ready before the Python client connects. The default scene provides the visible floor used by the puck experiment.
- Open Windows search by pressing the Windows key.
- Type CoppeliaSim and press Enter.
You should see CoppeliaSim's default scene with a floor in the main view. Keep this scene open for the connection check.
- Use the simulation controls in the toolbar to stop the simulation if it is running.
The simulator should now remain stopped. This gives the Python script control over when simulation time begins.
- Open the physics-engine selector in the CoppeliaSim toolbar.
- Select MuJoCo.
The physics-engine selector should show MuJoCo. The default floor should remain visible while the simulation stays stopped.
Create and run the connection check
The connection script asks CoppeliaSim for its product version before enabling synchronized stepping. It starts the simulation and advances exactly one step before stopping cleanly.
- Return to Notepad from earlier.
- Create a new document by pressing Ctrl+N.
- Add the connection check by pasting this code:
import time
from coppeliasim_zmqremoteapi_client import RemoteAPIClient
client = RemoteAPIClient()
sim = client.require("sim")
version = sim.getStringProperty(sim.handle_app, "productVersion")
print(f"Connected to CoppeliaSim {version}")
sim.setStepping(True)
sim.startSimulation()
sim.step()
print(f"Simulation time: {sim.getSimulationTime():.3f} seconds")
sim.stopSimulation()
while sim.getSimulationState() != sim.simulation_stopped:
time.sleep(0.05)
What does this code do?
- The RemoteAPIClient connection gives Python access to CoppeliaSim's sim API.
- The productVersion property confirms which CoppeliaSim release accepted the connection.
- Synchronized stepping makes sim.step() advance one controlled simulation step.
- The final loop waits until CoppeliaSim reports the stopped state before the script exits.
You should now see the complete connection script in Notepad. Its final indented line belongs inside the while loop.
- Open Notepad's save dialog by pressing Ctrl+S.
- Save the document as check_setup.py inside the local project folder.
The Notepad tab should now show check_setup.py. CoppeliaSim should remain open on the stopped default scene.
Before you run the check, make a quick prediction about which printed line will prove that Python advanced the simulator.
- Return to the active PowerShell window.
- Run the connection check by running this command:
python check_setup.py
What should you see?
The first line starts with Connected to CoppeliaSim and reports version 4.10.0. This proves the ZeroMQ Remote API connection succeeded.
The next line starts with Simulation time: and shows a value greater than zero. This proves Python advanced one synchronized simulation step.
That connection is the foundation locked in. Python can now control CoppeliaSim one measured step at a time.
Connection check not completing?
- Confirm CoppeliaSim is open on its default scene while the script runs.
- Confirm the simulation is stopped before retrying the script.
- Confirm the physics-engine selector shows MuJoCo.
- Compare check_setup.py with the full reference below.
- If the connection reports an older CoppeliaSim version, update to 4.10.0 from the official CoppeliaSim website.
- Still stuck? Help me debug the CoppeliaSim connection.
✔️ Awesome, I've got everything!
Great. Double check that requirements.txt and check_setup.py are saved inside your project folder.
ⓧ I'd like to double check the full code
Compare your dependency file with this complete reference.
coppeliasim-zmqremoteapi-client==2.0.4
torch==2.14.1
What should match?
Both package names and version pins must match exactly. Keep each dependency on its own line.
Compare your connection script with this complete reference.
import time
from coppeliasim_zmqremoteapi_client import RemoteAPIClient
client = RemoteAPIClient()
sim = client.require("sim")
version = sim.getStringProperty(sim.handle_app, "productVersion")
print(f"Connected to CoppeliaSim {version}")
sim.setStepping(True)
sim.startSimulation()
sim.step()
print(f"Simulation time: {sim.getSimulationTime():.3f} seconds")
sim.stopSimulation()
while sim.getSimulationState() != sim.simulation_stopped:
time.sleep(0.05)
What should match?
Check every function name and string against the reference. Keep the final time.sleep(0.05) line indented inside the loop.
Your synchronized simulator bridge is working. Next, you will use it to create a puck and collect the first state-action-next-state transitions.
Collect Biased Puck Data
Your Python connection now controls CoppeliaSim one synchronized step at a time. A world model needs repeatable transitions before it can learn how an action changes motion.
This step turns puck movement into state-action-next-state transitions. You will deliberately limit the training actions to positive X forces so the dataset hides leftward motion and sideways motion.
In this step, get ready to:
- Build a synchronized puck data collector.
- Collect 360 transitions using positive X forces.
- Collect 120 test transitions using unseen force directions.
Build the synchronized collector
The collector creates a temporary dynamic puck under the MuJoCo physics engine. Each synchronized step records the current state and force before calculating the resulting state change.
- Press the Windows key to open Windows search.
- Type Windows Notepad into the search field.
- Press Enter to open Windows Notepad.
- Create a blank file for collect_data.py.
- Paste the following code sections into the file in the order shown:
import argparse
import random
import time
import torch
from coppeliasim_zmqremoteapi_client import RemoteAPIClient
def wait_for_stop(sim):
while sim.getSimulationState() != sim.simulation_stopped:
time.sleep(0.05)
def create_puck(sim):
puck = sim.createPrimitiveShape(
sim.primitiveshape_cuboid,
[0.12, 0.12, 0.04],
)
sim.setBoolProperty(puck, "dynamic", True)
sim.setBoolProperty(puck, "respondable", True)
sim.setFloatProperty(puck, "mass", 1.0)
sim.setObjectPosition(puck, [0.0, 0.0, 0.08])
return puck
How does the puck get created?
- The imports provide command-line parsing and deterministic random actions. They also provide timing and tensor storage.
- The wait_for_stop() helper waits until CoppeliaSim confirms that a simulation has stopped.
- The create_puck() helper creates a cuboid with dimensions of 0.12 by 0.12 by 0.04.
- The puck is dynamic and respondable. Its mass is set to 1.0 kilogram so forces can move it across the floor.
def read_state(sim, puck):
position = sim.getObjectPosition(puck)
linear_velocity, _ = sim.getVelocity(puck)
return [
position[0],
position[1],
linear_velocity[0],
linear_velocity[1],
]
def sample_action(rng, policy):
if policy == "biased":
return [rng.uniform(0.10, 0.35), 0.0]
if policy == "test":
return [rng.uniform(-0.35, 0.05), rng.uniform(-0.35, 0.35)]
return [rng.uniform(-0.35, 0.35), rng.uniform(-0.35, 0.35)]
How are states and actions represented?
- The read_state() helper returns the puck's X position and Y position. It also returns its X velocity and Y velocity.
- The biased policy samples X forces between 0.10 and 0.35. Its Y force always remains 0.0.
- The test policy includes X forces from -0.35 to 0.05. Its Y forces span -0.35 to 0.35.
def collect(policy, episodes, output, seed):
rng = random.Random(seed)
client = RemoteAPIClient()
sim = client.require("sim")
sim.setStepping(True)
puck = create_puck(sim)
inputs = []
targets = []
What starts each collection run?
- The fixed random seed makes the chosen actions repeatable across runs.
- The remote client connects to the open CoppeliaSim session. Stepping mode gives Python control over each simulation step.
- The inputs list stores states plus actions. The targets list stores the state changes caused by those actions.
try:
for episode in range(episodes):
start_position = [
rng.uniform(-0.20, 0.20),
rng.uniform(-0.20, 0.20),
0.08,
]
sim.setObjectPosition(puck, start_position)
sim.setVector3Property(puck, "initLinearVelocity", [0.0, 0.0, 0.0])
sim.startSimulation()
for _ in range(10):
sim.step()
state = read_state(sim, puck)
action = [0.0, 0.0]
How does each episode begin?
- Each episode places the puck near the center of the floor at a randomized X and Y position.
- The initial linear velocity is reset to zero so the previous episode cannot affect the next one.
- Ten synchronized steps let the puck settle before the first state is recorded.
for step in range(60):
if step % 6 == 0:
action = sample_action(rng, policy)
sim.addForce(
puck,
[0.0, 0.0, 0.0],
[action[0], action[1], 0.0],
)
sim.step()
next_state = read_state(sim, puck)
delta = [
next_state[index] - state[index]
for index in range(4)
]
inputs.append(state + action)
targets.append(delta)
state = next_state
How does one transition get recorded?
- A new force action is sampled every six steps. This lets the puck respond to each force across several observations.
- The force is applied at the puck's center with a zero Z component. This keeps the experiment focused on motion across the floor.
- The input contains [x, y, vx, vy, fx, fy]. The target contains [Δx, Δy, Δvx, Δvy].
sim.stopSimulation()
wait_for_stop(sim)
print(f"Episode {episode + 1}/{episodes} complete")
finally:
if sim.getSimulationState() != sim.simulation_stopped:
sim.stopSimulation()
wait_for_stop(sim)
sim.removeObject(puck)
torch.save(
{
"inputs": torch.tensor(inputs),
"targets": torch.tensor(targets),
},
output,
)
print(f"Saved {len(inputs)} transitions to {output}")
How does the collector finish safely?
- The simulation stops after every episode. The collector waits for CoppeliaSim to confirm that stopped state.
- The finally block stops an unfinished simulation if the run exits early. It also removes the temporary puck.
- PyTorch saves one dictionary containing the inputs tensor and the targets tensor.
def main():
parser = argparse.ArgumentParser()
parser.add_argument(
"--policy",
choices=["biased", "balanced", "test"],
required=True,
)
parser.add_argument("--episodes", type=int, required=True)
parser.add_argument("--output", required=True)
parser.add_argument("--seed", type=int, default=7)
args = parser.parse_args()
collect(args.policy, args.episodes, args.output, args.seed)
if __name__ == "__main__":
main()
How do the command options work?
- The --policy option selects the biased policy or one of the wider action policies.
- The --episodes option controls how many 60-step episodes the collector records.
- The --output option names the saved dataset. The default seed keeps repeated experiments comparable.
- Save the file as collect_data.py inside the Windows project folder that contains check_setup.py.
- Close Windows Notepad.
- Return to the activated .venv session in PowerShell.
✔️ Awesome, I've got everything!
Your collector is saved. It can create the puck and capture synchronized transitions before removing the temporary simulation object.
ⓧ I'd like to double check the full code
- Compare collect_data.py with this complete file:
import argparse
import random
import time
import torch
from coppeliasim_zmqremoteapi_client import RemoteAPIClient
def wait_for_stop(sim):
while sim.getSimulationState() != sim.simulation_stopped:
time.sleep(0.05)
def create_puck(sim):
puck = sim.createPrimitiveShape(
sim.primitiveshape_cuboid,
[0.12, 0.12, 0.04],
)
sim.setBoolProperty(puck, "dynamic", True)
sim.setBoolProperty(puck, "respondable", True)
sim.setFloatProperty(puck, "mass", 1.0)
sim.setObjectPosition(puck, [0.0, 0.0, 0.08])
return puck
def read_state(sim, puck):
position = sim.getObjectPosition(puck)
linear_velocity, _ = sim.getVelocity(puck)
return [
position[0],
position[1],
linear_velocity[0],
linear_velocity[1],
]
def sample_action(rng, policy):
if policy == "biased":
return [rng.uniform(0.10, 0.35), 0.0]
if policy == "test":
return [rng.uniform(-0.35, 0.05), rng.uniform(-0.35, 0.35)]
return [rng.uniform(-0.35, 0.35), rng.uniform(-0.35, 0.35)]
def collect(policy, episodes, output, seed):
rng = random.Random(seed)
client = RemoteAPIClient()
sim = client.require("sim")
sim.setStepping(True)
puck = create_puck(sim)
inputs = []
targets = []
try:
for episode in range(episodes):
start_position = [
rng.uniform(-0.20, 0.20),
rng.uniform(-0.20, 0.20),
0.08,
]
sim.setObjectPosition(puck, start_position)
sim.setVector3Property(puck, "initLinearVelocity", [0.0, 0.0, 0.0])
sim.startSimulation()
for _ in range(10):
sim.step()
state = read_state(sim, puck)
action = [0.0, 0.0]
for step in range(60):
if step % 6 == 0:
action = sample_action(rng, policy)
sim.addForce(
puck,
[0.0, 0.0, 0.0],
[action[0], action[1], 0.0],
)
sim.step()
next_state = read_state(sim, puck)
delta = [
next_state[index] - state[index]
for index in range(4)
]
inputs.append(state + action)
targets.append(delta)
state = next_state
sim.stopSimulation()
wait_for_stop(sim)
print(f"Episode {episode + 1}/{episodes} complete")
finally:
if sim.getSimulationState() != sim.simulation_stopped:
sim.stopSimulation()
wait_for_stop(sim)
sim.removeObject(puck)
torch.save(
{
"inputs": torch.tensor(inputs),
"targets": torch.tensor(targets),
},
output,
)
print(f"Saved {len(inputs)} transitions to {output}")
def main():
parser = argparse.ArgumentParser()
parser.add_argument(
"--policy",
choices=["biased", "balanced", "test"],
required=True,
)
parser.add_argument("--episodes", type=int, required=True)
parser.add_argument("--output", required=True)
parser.add_argument("--seed", type=int, default=7)
args = parser.parse_args()
collect(args.policy, args.episodes, args.output, args.seed)
if __name__ == "__main__":
main()
What should match?
The complete file defines wait_for_stop(), create_puck(), read_state(), sample_action(), collect(), and main(). Their order and indentation must match the reference.
Collector file not saving correctly?
- Check that the filename is exactly collect_data.py. Remove any extra .txt extension added by Notepad.
- Check that the file sits beside check_setup.py in the same Windows project folder.
- Still stuck? Help me check my collect_data.py file against the required CoppeliaSim collector structure.
Collect the biased training set
The first training dataset contains only positive X forces. This action coverage lets the puck demonstrate an immediate visible response while withholding leftward and sideways behavior from the future model.
- Arrange PowerShell and CoppeliaSim so both remain visible.
- Confirm that the simulation is stopped in CoppeliaSim.
Before you run the collector, which directions do you expect the puck to explore when every Y force is zero?
- Collect six biased episodes by running this command:
python collect_data.py --policy biased --episodes 6 --output biased.pt
What should you see?
You will see the temporary puck move mainly along the positive X direction in CoppeliaSim. PowerShell prints progress for six episodes before printing Saved 360 transitions to biased.pt.
The saved inputs tensor has shape [360, 6]. The saved targets tensor has shape [360, 4].
That first dataset is ready. You now have 360 synchronized examples of how the puck responds to a deliberately narrow action policy.
Puck not moving or dataset not saved?
- Confirm that CoppeliaSim remains open with MuJoCo selected.
- Confirm that PowerShell is still inside the Windows project folder containing collect_data.py.
- Still stuck? Help me troubleshoot my biased CoppeliaSim data collection run.
Collect the held-out test set
A useful test set asks the collector for behavior that the training set omitted. This policy introduces negative X forces plus both Y directions while leaving the biased training data unchanged.
Before you run this final check, how do you expect the puck's path to differ from the positive-X-only run?
- Collect two held-out test episodes by running this command:
python collect_data.py --policy test --episodes 2 --output test.pt
What should you see?
You will see the puck explore negative X movement and both Y directions. PowerShell prints progress for two episodes before printing Saved 120 transitions to test.pt.
The saved inputs tensor has shape [120, 6]. The saved targets tensor has shape [120, 4].
Test collection not completing?
- Check that the first collection run finished before starting the test collection.
- Confirm that CoppeliaSim shows a stopped simulation after the command exits. The temporary puck should also be removed.
- Still stuck? Help me troubleshoot my test.pt collection run.
Your datasets now create a deliberate action-coverage gap. Next, you will train the first world model and see whether low training error survives those unseen directions.
Train the Biased World Model
Your CoppeliaSim collector has turned puck motion into 360 biased training transitions and 120 test transitions. Those files now let you check whether a world model can handle actions beyond its training experience.
A low training error can look convincing when important action directions are absent. The held-out test set gives you a direct way to measure this distribution shift after training the model with PyTorch.
In this step, get ready to:
- Define a neural network that predicts four state changes from six state-action values.
- Create a training script that measures performance on both datasets.
- Train the biased model and document its action-coverage failure.
Define the action-conditioned network
An action-conditioned model receives the puck's four current state values plus its two force values. It predicts the four changes that produce the next state.
- Switch back to the PowerShell window with the activated .venv.
- Select Yes if Windows Notepad asks whether to create the new file.
- Create world_model.py in the current project folder by running this command:
notepad world_model.py
What does this command do?
The command asks Windows Notepad to open world_model.py from PowerShell's current folder. Saving the new document places the file beside your existing project files.
- Paste this model definition into the blank world_model.py document:
from torch import nn
def build_model():
return nn.Sequential(
nn.Linear(6, 64),
nn.ReLU(),
nn.Linear(64, 64),
nn.ReLU(),
nn.Linear(64, 4),
)
What does this model do?
- The first nn.Linear(6, 64) layer maps four state values plus two force values into 64 learned features.
- The two nn.ReLU() layers let the network represent nonlinear motion.
- The final nn.Linear(64, 4) layer predicts changes in position and velocity.
- Select File from the Notepad menu bar.
- Select Save.
- Check the Notepad title bar. You should see world_model.py as the saved filename.
✔️ Awesome, I've got everything!
Your world_model.py file now defines the complete six-input to four-output network.
ⓧ I'd like to double check the full code
from torch import nn
def build_model():
return nn.Sequential(
nn.Linear(6, 64),
nn.ReLU(),
nn.Linear(64, 64),
nn.ReLU(),
nn.Linear(64, 4),
)
What should match?
Compare every layer size and activation with your saved file. The input width is six because each transition contains four state values plus two force values.
Model file not saved?
- Check that the Notepad title bar shows world_model.py without a .txt suffix.
- Return to the PowerShell window from earlier before opening the file again.
- Still stuck? Help me check why world_model.py is missing or has the wrong filename.
Build the training loop
The trainer needs to load both datasets through the same path before it can compare them fairly. It also needs one shared loss function so each reported error measures the same four predicted state changes.
- Switch back to the PowerShell window from earlier.
- Select Yes if Notepad asks whether to create the new file.
- Create train_model.py in the current project folder by running this command:
notepad train_model.py
What does this command do?
The command opens a document named train_model.py in the project folder. This keeps the trainer beside world_model.py and the two dataset files it imports.
- Paste this first section into the blank train_model.py document:
import argparse
import math
import torch
from torch import nn
from world_model import build_model
def load_dataset(path):
payload = torch.load(path, weights_only=True)
return payload["inputs"], payload["targets"]
def main():
parser = argparse.ArgumentParser()
parser.add_argument("--train", required=True)
parser.add_argument("--test", required=True)
parser.add_argument("--model", required=True)
parser.add_argument("--epochs", type=int, default=800)
args = parser.parse_args()
train_inputs, train_targets = load_dataset(args.train)
test_inputs, test_targets = load_dataset(args.test)
model = build_model()
loss_function = nn.MSELoss()
optimizer = torch.optim.Adam(model.parameters(), lr=0.01)
What does this setup do?
- The load_dataset() function retrieves the saved input and target tensors with weights_only=True.
- The command-line arguments choose the training dataset and test dataset. They also choose the checkpoint filename and epoch count.
- The trainer creates the shared network with build_model().
- The nn.MSELoss() function measures prediction error. Adam updates the model with a learning rate of 0.01.
- Select File from the Notepad menu bar.
- Select Save.
- Check the Notepad title bar. You should see train_model.py as the saved filename.
The model and datasets are now connected inside the trainer. The remaining section performs optimization before calculating both final errors.
- Place the cursor directly below the existing optimizer line in train_model.py.
- Paste the remaining training and evaluation logic:
model.train()
for epoch in range(args.epochs):
prediction = model(train_inputs)
loss = loss_function(prediction, train_targets)
optimizer.zero_grad()
loss.backward()
optimizer.step()
if (epoch + 1) % 200 == 0:
print(f"Epoch {epoch + 1}: MSE={loss.item():.8f}")
model.eval()
with torch.no_grad():
train_mse = loss_function(model(train_inputs), train_targets).item()
test_mse = loss_function(model(test_inputs), test_targets).item()
train_rmse = math.sqrt(train_mse)
test_rmse = math.sqrt(test_mse)
torch.save(model.state_dict(), args.model)
print(f"Training RMSE: {train_rmse:.6f}")
print(f"Balanced-test RMSE: {test_rmse:.6f}")
print(f"Saved model state to {args.model}")
if __name__ == "__main__":
main()
How does training work?
- Each epoch predicts every training target before calculating mean squared error.
- The optimizer clears old gradients before backpropagation updates the network.
- Evaluation disables gradient tracking before calculating training and test errors.
- The script converts both errors to root mean squared error before saving the model's state dictionary.
- Select File from the Notepad menu bar.
- Select Save.
✔️ Awesome, I've got everything!
Your trainer now loads both datasets and saves a checkpoint after 800 epochs.
ⓧ I'd like to double check the full code
import argparse
import math
import torch
from torch import nn
from world_model import build_model
def load_dataset(path):
payload = torch.load(path, weights_only=True)
return payload["inputs"], payload["targets"]
def main():
parser = argparse.ArgumentParser()
parser.add_argument("--train", required=True)
parser.add_argument("--test", required=True)
parser.add_argument("--model", required=True)
parser.add_argument("--epochs", type=int, default=800)
args = parser.parse_args()
train_inputs, train_targets = load_dataset(args.train)
test_inputs, test_targets = load_dataset(args.test)
model = build_model()
loss_function = nn.MSELoss()
optimizer = torch.optim.Adam(model.parameters(), lr=0.01)
model.train()
for epoch in range(args.epochs):
prediction = model(train_inputs)
loss = loss_function(prediction, train_targets)
optimizer.zero_grad()
loss.backward()
optimizer.step()
if (epoch + 1) % 200 == 0:
print(f"Epoch {epoch + 1}: MSE={loss.item():.8f}")
model.eval()
with torch.no_grad():
train_mse = loss_function(model(train_inputs), train_targets).item()
test_mse = loss_function(model(test_inputs), test_targets).item()
train_rmse = math.sqrt(train_mse)
test_rmse = math.sqrt(test_mse)
torch.save(model.state_dict(), args.model)
print(f"Training RMSE: {train_rmse:.6f}")
print(f"Balanced-test RMSE: {test_rmse:.6f}")
print(f"Saved model state to {args.model}")
if __name__ == "__main__":
main()
What should match?
Compare the argument names and training order with your saved file. The trainer must evaluate both datasets before saving model.state_dict().
Trainer file looks incomplete?
- Check that the second section begins inside main() with four spaces before model.train().
- Confirm that world_model.py sits in the same project folder as train_model.py.
- Still stuck? Help me compare my training script with the expected PyTorch structure.
Train the biased model and record the gap
The trainer now faces two different action distributions. Before you run it, do you expect the training and balanced-test errors to stay close together?
- Switch back to the PowerShell window with the activated .venv.
- Train on biased.pt and evaluate on test.pt by running this command:
python train_model.py --train biased.pt --test test.pt --model biased_model.pt
What should I see?
You should see mean squared error updates at regular epoch intervals. The final lines show Training RMSE and Balanced-test RMSE.
The balanced-test value should be higher than the training value. You should also see confirmation that the model state was saved to biased_model.pt.
Good work. biased_model.pt now holds your first trained world model.
Why did the test error rise?
The training data contains positive X forces with zero Y force. The test data introduces negative X forces and nonzero Y forces.
The lower training error shows that optimization fitted the sampled transitions. The larger test error points to missing action coverage instead of an observed training optimization failure.
- Record the printed training value here: your training RMSE.
- Record the printed balanced-test value here: your balanced-test RMSE.
- Select Yes if Notepad asks whether to create the notes file.
- Create experiment_notes.txt in the current project folder by running this command:
notepad experiment_notes.txt
What does this command do?
The command opens a project note named experiment_notes.txt in Notepad. This file keeps the measured errors beside the checkpoint they describe.
- Type Training RMSE: followed by your training RMSE.
- Type Balanced-test RMSE: followed by your balanced-test RMSE on the next line.
- Add a sentence explaining that training used only positive X forces with zero Y force.
- Add a sentence explaining that the test data contains unseen negative X and nonzero Y directions.
- State that the larger test error indicates missing action coverage instead of an observed training optimization failure.
- Select File from the Notepad menu bar.
- Select Save.
That is the deliberate shortfall captured. Your checkpoint fits the biased experience while the held-out actions reveal what it has not learned.
You have measured the model's action-coverage gap without changing its architecture. Next, you will test whether broader experience improves the same network.
Train with Balanced Actions
The biased checkpoint exposed a distribution shift. Its training data covered rightward force while the test set introduced unseen directions.
Balanced exploration now covers both axes in both directions. The PyTorch world model stays unchanged.
Any improvement therefore points to the training data.
In this step, get ready to:
- Collect 480 transitions using balanced two-axis forces.
- Train the unchanged world model on the balanced dataset.
- Document the controlled comparison between both checkpoints.
Collect balanced action data
The existing CoppeliaSim collector already includes a balanced policy. It samples X-axis force independently from Y-axis force.
MuJoCo remains selected. This keeps the physics environment fixed.
- Keep CoppeliaSim visible beside PowerShell while the command runs.
- Use the existing balanced policy for eight episodes by running this command:
python collect_data.py --policy balanced --episodes 8 --output balanced.pt
How does this rebalance the data?
- The balanced policy samples symmetric forces along the X axis.
- The same policy samples symmetric forces along the Y axis.
- Eight 60-step episodes produce 480 state-action-next-state transitions.
- Each input still contains four state values plus two force values.
- Check the final line in PowerShell.
PowerShell should end with Saved 480 transitions to balanced.pt. That gives your world model experience with force in every planned direction.
Did the collection stop early?
- A simulation that was already running can block the collector. Stop it before your next attempt.
- An unexpected puck response can indicate a different physics engine. Confirm the physics-engine toolbar selector still shows MuJoCo.
- Still stuck? Help me troubleshoot balanced puck data collection.
Train the unchanged model
The network architecture stays fixed at 6-64-64-4. The existing trainer also keeps its optimizer settings fixed.
Before you train, do you expect broader action coverage to reduce the error on the unchanged test set?
- Use balanced.pt as the training data with test.pt held fixed by running this command:
python train_model.py --train balanced.pt --test test.pt --model balanced_model.pt
What stays controlled?
- The --train argument changes the training source to balanced.pt.
- The --test argument keeps test.pt as the shared evaluation dataset.
- The trainer still builds the model through world_model.build_model().
- The --model argument saves the new state dictionary as balanced_model.pt.
- Check the final lines in PowerShell.
- Record the printed training RMSE here: your balanced training RMSE.
- Record the printed balanced-test RMSE here: your balanced-test RMSE.
PowerShell should print Saved model state to balanced_model.pt. The balanced-test RMSE should be lower than the biased model's balanced-test RMSE.
You have now improved generalization without adding a single layer to the model.
Is the balanced-test RMSE still higher?
- Confirm the earlier collection ended with 480 transitions in balanced.pt.
- Confirm the training command used balanced.pt for training.
- Still stuck? Help me investigate the balanced model result.
Record and verify the comparison
Your experiment notes turn two terminal results into a controlled conclusion. The evidence needs to separate the data change from the fixed model design.
What should the note capture?
- The balanced model's training RMSE is your balanced training RMSE.
- The balanced model's balanced-test RMSE is your balanced-test RMSE.
- Symmetric two-axis action coverage reduced the balanced-test error.
- The 6-64-64-4 network architecture stayed fixed.
- In experiment_notes.txt from earlier, add one balanced-model entry using those four details.
- Save experiment_notes.txt.
Before you make the final comparison, which dataset do you expect to generalize better across the held-out action directions?
- Compare the balanced-test RMSE in experiment_notes.txt with the biased model's balanced-test RMSE.
- Confirm the final training output includes Saved model state to balanced_model.pt.
- Confirm the earlier collection output includes Saved 480 transitions to balanced.pt.
You should see a lower balanced-test RMSE for the balanced model. The two saved-file messages confirm that balanced.pt and balanced_model.pt now exist.
That's the data-first debugging result locked in: broader action coverage fixed the model's blind spot while its capacity stayed constant.
Your model now has balanced experience across both motion axes. Next, you'll test how its one-step predictions behave when they are chained into an imagined trajectory.
Compare Real and Imagined Rollouts
Your balanced model now handles held-out actions more accurately inside CoppeliaSim. That one-step result leaves the planning problem unanswered. Can the model stay useful when every prediction becomes part of its next input?
A recursive rollout feeds each predicted state back into the PyTorch model. In this step, you'll compare that imagined path with an 80-step MuJoCo path.
In this step, get ready to:
- Create the recursive rollout evaluator.
- Generate a browser report comparing real motion with imagined motion.
- Measure recursive position RMSE and explain where trajectory drift begins.
Build the rollout evaluator
The evaluator keeps one trajectory from the simulator. It keeps a second trajectory whose predicted state becomes the model's next input.
This recursive path reveals errors that a one-step test can hide. The script also turns both paths into a browser-readable report.
- Press the Windows key to open Windows search.
- Type Notepad into the search field.
- Press Enter to open Notepad.
- Select the second tab below to reveal the complete evaluator.
- Copy all the code shown in that tab.
- Paste the code into the blank Notepad document.
✔️ Awesome, I've got everything!
Your evaluator code is ready. Continue below to save it as evaluate_rollout.py.
ⓧ I'd like to double check the full code
import argparse
import math
import torch
from coppeliasim_zmqremoteapi_client import RemoteAPIClient
from collect_data import create_puck, read_state, wait_for_stop
from world_model import build_model
def scale_points(paths, width=640, height=420, margin=35):
all_x = [point[0] for path in paths for point in path]
all_y = [point[1] for path in paths for point in path]
x_min, x_max = min(all_x), max(all_x)
y_min, y_max = min(all_y), max(all_y)
x_span = max(x_max - x_min, 0.001)
y_span = max(y_max - y_min, 0.001)
scaled_paths = []
for path in paths:
scaled = []
for x_value, y_value in path:
x_pixel = margin + (x_value - x_min) * (width - 2 * margin) / x_span
y_pixel = height - margin - (
(y_value - y_min) * (height - 2 * margin) / y_span
)
scaled.append((x_pixel, y_pixel))
scaled_paths.append(scaled)
return scaled_paths
def polyline(points):
return " ".join(f"{x_value:.1f},{y_value:.1f}" for x_value, y_value in points)
def write_report(path, actual_xy, predicted_xy, rmse):
actual_scaled, predicted_scaled = scale_points([actual_xy, predicted_xy])
html = f"""<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8">
<title>CoppeliaSim World Model Rollout</title>
<style>
body {{ font-family: Arial, sans-serif; max-width: 760px; margin: 30px auto; }}
svg {{ background: #f7f7f7; border: 1px solid #cccccc; }}
.legend {{ display: flex; gap: 24px; margin: 12px 0; }}
.blue {{ color: #1261a0; }}
.orange {{ color: #d46b08; }}
</style>
</head>
<body>
<h1>Real vs imagined rollout</h1>
<p>Recursive position RMSE: {rmse:.6f}</p>
<div class="legend">
<span class="blue">Blue: CoppeliaSim</span>
<span class="orange">Orange: world model</span>
</div>
<svg width="640" height="420" viewBox="0 0 640 420">
<polyline points="{polyline(actual_scaled)}" fill="none" stroke="#1261a0" stroke-width="3" />
<polyline points="{polyline(predicted_scaled)}" fill="none" stroke="#d46b08" stroke-width="3" />
</svg>
</body>
</html>
"""
with open(path, "w", encoding="utf-8") as report_file:
report_file.write(html)
def main():
parser = argparse.ArgumentParser()
parser.add_argument("--model", required=True)
parser.add_argument("--report", required=True)
args = parser.parse_args()
model = build_model()
model.load_state_dict(torch.load(args.model, weights_only=True))
model.eval()
client = RemoteAPIClient()
sim = client.require("sim")
sim.setStepping(True)
puck = create_puck(sim)
actual_states = []
predicted_states = []
try:
sim.setObjectPosition(puck, [0.0, 0.0, 0.08])
sim.setVector3Property(puck, "initLinearVelocity", [0.0, 0.0, 0.0])
sim.startSimulation()
for _ in range(10):
sim.step()
actual_state = read_state(sim, puck)
predicted_state = actual_state.copy()
actual_states.append(actual_state)
predicted_states.append(predicted_state.copy())
for step in range(80):
action = [
0.30 * math.sin(step / 8.0),
0.30 * math.cos(step / 11.0),
]
with torch.no_grad():
model_input = torch.tensor([predicted_state + action])
predicted_delta = model(model_input)[0]
delta_values = [predicted_delta[index].item() for index in range(4)]
predicted_state = [
predicted_state[index] + delta_values[index]
for index in range(4)
]
sim.addForce(
puck,
[0.0, 0.0, 0.0],
[action[0], action[1], 0.0],
)
sim.step()
actual_state = read_state(sim, puck)
predicted_states.append(predicted_state.copy())
actual_states.append(actual_state)
sim.stopSimulation()
wait_for_stop(sim)
finally:
if sim.getSimulationState() != sim.simulation_stopped:
sim.stopSimulation()
wait_for_stop(sim)
sim.removeObject(puck)
squared_error = 0.0
for actual, predicted in zip(actual_states, predicted_states):
squared_error += (actual[0] - predicted[0]) ** 2
squared_error += (actual[1] - predicted[1]) ** 2
rmse = math.sqrt(squared_error / (2 * len(actual_states)))
actual_xy = [(state[0], state[1]) for state in actual_states]
predicted_xy = [(state[0], state[1]) for state in predicted_states]
write_report(args.report, actual_xy, predicted_xy, rmse)
print(f"Recursive position RMSE: {rmse:.6f}")
print(f"Saved rollout report to {args.report}")
if __name__ == "__main__":
main()
What does this evaluator do?
- The scale_points() function maps both trajectories into one 640 by 420 plotting area.
- The polyline() function converts coordinates into points for the SVG paths.
- The write_report() function embeds both paths into rollout_report.html with the position error.
- The X force uses 0.30 * math.sin(step / 8.0). The Y force uses 0.30 * math.cos(step / 11.0).
- The predicted trajectory updates predicted_state from its own four-value deltas. Simulator states never correct this imagined trajectory.
- The cleanup block stops the simulation. It also removes the temporary puck.
- Press Ctrl+Shift+S to open the Save As dialog.
- Select your project folder from earlier.
- Enter evaluate_rollout.py in the File name field.
- Select All Files (*.*) from the Save as type list.
- Click Save.
- Check the Notepad title bar.
You should see evaluate_rollout.py in the title bar. This confirms the script has the Python extension.
Seeing a text-file extension?
- Open the Save As dialog again with Ctrl+Shift+S.
- Select All Files (*.*) before saving evaluate_rollout.py again.
- Still stuck? Help me save a Python file from Notepad without a .txt extension.
Run the recursive comparison
The evaluator starts a temporary puck from one shared initial state. The simulator receives the same 80 actions as the model.
The real path reads every new state from MuJoCo. The imagined path advances entirely through predictions after its initial state.
Before you run the comparison, how closely do you expect the imagined path to follow MuJoCo after every prediction feeds the next input?
- Switch back to PowerShell from earlier.
- Run the real and imagined rollout comparison with this command:
python evaluate_rollout.py --model balanced_model.pt --report rollout_report.html
What does this command do?
The --model balanced_model.pt argument loads the balanced checkpoint. The --report rollout_report.html argument names the generated report.
The script performs 80 synchronized simulator steps. It calculates position error across the complete recursive trajectory.
You'll see Recursive position RMSE: followed by a measured value. The next line says Saved rollout report to rollout_report.html.
That completes the full evaluation loop. The temporary puck is removed while CoppeliaSim remains ready for another run.
Evaluation not completing?
- Confirm CoppeliaSim still shows MuJoCo while the simulation is stopped.
- Confirm the PowerShell prompt still shows the activated .venv environment.
- Confirm balanced_model.pt sits in the same project folder as evaluate_rollout.py.
- Still stuck? Help me troubleshoot my recursive rollout evaluation.
Before you open the report, where do you expect the orange trajectory to begin separating from the blue trajectory?
- Press the Windows key to open Windows search.
- Type File Explorer into the search field.
- Press Enter to open File Explorer.
- Select your project folder from earlier.
- Double-click rollout_report.html.
Your browser shows an HTML report titled CoppeliaSim World Model Rollout. You should see a blue CoppeliaSim path beside an orange world-model path.
The report also shows a recursive position RMSE. You now have a visual measurement of how prediction error compounds across an imagined trajectory.
- Record the displayed MuJoCo recursive position RMSE here: your MuJoCo recursive rollout RMSE.
Explain where imagination drifts
A single RMSE value summarizes the whole rollout. The trajectory shape reveals when the model's recursive inputs start pulling it away from the simulator.
Look for the first sustained gap between the blue path and the orange path. Later separation reflects earlier prediction errors flowing into future predictions.
- Return to the File Explorer window from earlier.
- Double-click experiment_notes.txt.
- Add the measured MuJoCo recursive rollout RMSE of your MuJoCo recursive rollout RMSE.
- Add one sentence describing where the orange imagined path begins to drift from the blue simulator path.
- Save experiment_notes.txt.
Before you rerun the evaluator, do you expect the two trajectories to share their starting point after a fresh simulation?
- Switch back to PowerShell from earlier.
- Regenerate the report to confirm the complete evaluation pipeline with this command:
python evaluate_rollout.py --model balanced_model.pt --report rollout_report.html
Why run the evaluation again?
This command recreates the temporary puck from the same initial state. It regenerates rollout_report.html from the fixed balanced checkpoint.
A successful second run confirms that evaluation includes its own simulator setup. It also confirms that cleanup leaves CoppeliaSim ready for reuse.
PowerShell again prints the recursive position RMSE. It also confirms that rollout_report.html was saved.
- Return to the browser tab from earlier.
- Refresh rollout_report.html.
- Compare the blue path with the orange path from their shared starting point.
You should see both paths start from the same position. Any later separation shows recursive prediction error accumulating across the 80-step rollout.
Report not refreshing?
- Confirm the browser address ends with rollout_report.html from your project folder.
- Refresh the browser after PowerShell prints the saved-report confirmation.
- Still stuck? Help me verify that my regenerated rollout report is opening from the correct project folder.
You have now measured more than one-step accuracy. Your report shows how the balanced world model behaves when it must rely on its own imagined future.
Secret mission
Test Sim-to-Sim Transfer
Keep your trained model fixed while changing CoppeliaSim from MuJoCo to Bullet Physics. You will measure how a new simulator changes recursive rollout accuracy without retraining.
Clean Up Your Resources
Clean Up Your Resources
Your local CoppeliaSim experiment has no ongoing project usage charges. Decide whether to keep the project folder, pause your session, or delete the folder.
Resources you used:
- The Python virtual environment stored in the .venv folder.
- The project scripts check_setup.py, collect_data.py, world_model.py, train_model.py, evaluate_rollout.py, and requirements.txt.
- The PyTorch datasets and checkpoints biased.pt, test.pt, balanced.pt, biased_model.pt, and balanced_model.pt.
- The experiment record in experiment_notes.txt plus the HTML reports rollout_report.html and bullet_report.html.
Keep everything running
Keep this setup if you want to repeat the experiment or extend the world model. No cleanup action is required.
- Keep the project folder so every script, dataset, checkpoint, note, and report remains available.
- Keep the current PowerShell session open so the .venv environment remains active.
- Keep CoppeliaSim open if you plan to continue comparing physics engines. The simulation remains stopped with Bullet Physics selected.
Pause - I'll come back to this later
Pause your local session to close the tools while preserving every experiment artifact. You can return to the same project folder later.
- Close CoppeliaSim from its application window.
- Switch back to the PowerShell session from earlier.
- Deactivate the virtual environment by running:
deactivate
What does this command do?
The deactivate command returns PowerShell to its normal Python environment. The .venv folder stays on disk.
- Close the PowerShell window.
Your experiment is safely paused. Every script, model, dataset, note, and report remains ready for your next session.
Delete - I don't want to use this again
Deleting the project folder permanently removes every dataset, checkpoint, note, report, script, and virtual-environment file. Your existing CoppeliaSim installation stays untouched.
- Close CoppeliaSim from its application window.
- Return to the activated PowerShell session from earlier.
- Deactivate the virtual environment by running:
deactivate
Why deactivate first?
The deactivate command releases the active virtual-environment session before you remove its files. This leaves PowerShell using its normal Python environment.
- Close the PowerShell window.
- Use Windows system search to open the system file browser.
- Locate the project folder using check_setup.py and bullet_report.html as identifiers.
- Permanently delete that project folder using the system file browser.
- Return to the location that previously held the project folder.
You should no longer see the project folder or any of its artifacts. Your local experiment is removed while your existing CoppeliaSim installation remains available.
Nice Work!
Nice Work!
Outstanding work! Your embodied AI pipeline now learns action-conditioned dynamics from synchronized CoppeliaSim transitions. The browser reports show how recursive predictions drift under MuJoCo and Bullet Physics.
You've learned how to:
- Build a synchronized data-collection loop through the ZeroMQ remote API for state-action-next-state transitions.
- Train an action-conditioned PyTorch world model to predict state deltas. Expose its out-of-distribution failure with biased actions. Fix the failure through balanced exploration while keeping the architecture unchanged.
- Generate an HTML trajectory report that overlays real simulator paths with recursively imagined paths. Measure the resulting drift with multi-step RMSE.
- Complete the Secret Mission by evaluating the unchanged balanced_model.pt checkpoint under Bullet Physics. Explain the RMSE difference as simulator distribution shift.
Ready to quiz yourself?