{ "cells": [ { "cell_type": "markdown", "id": "79a2bec1", "metadata": {}, "source": [ "# Simulator\n", "If we have a transition system, it might be nice to run a simulation. In this case, we have an MDP that models a hungry lion. Depending on the state it is in, it needs to decide whether it wants to 'rawr' or 'hunt' in order to prevent reaching the state 'dead'." ] }, { "cell_type": "code", "execution_count": 1, "id": "20a01918", "metadata": { "execution": { "iopub.execute_input": "2026-10-01T12:39:19.821593Z", "iopub.status.busy": "2026-10-01T12:39:19.821428Z", "iopub.status.idle": "2026-10-01T12:39:20.410436Z", "shell.execute_reply": "2026-10-01T12:39:20.409835Z" } }, "outputs": [ { "data": { "text/html": [ "\n", "\n", "\n", " \n", " Network\n", " \n", " \n", " \n", " \n", " \n", "
\n", " \n", " \n", " \n", " \n", "\n" ], "text/plain": [ "" ] }, "metadata": {}, "output_type": "display_data" }, { "data": { "text/plain": [ "" ] }, "execution_count": 1, "metadata": {}, "output_type": "execute_result" } ], "source": [ "from stormvogel import *\n", "import stormvogel\n", "\n", "lion = examples.create_lion_mdp()\n", "show(lion)" ] }, { "cell_type": "markdown", "id": "4ff8657d", "metadata": {}, "source": [ "Now, let's run a simulation of the lion! If we do not provide a scheduling function, then the simulator just does a random walk, taking a random choice each time." ] }, { "cell_type": "code", "execution_count": 2, "id": "c93cc571", "metadata": { "execution": { "iopub.execute_input": "2026-10-01T12:39:20.427963Z", "iopub.status.busy": "2026-10-01T12:39:20.427645Z", "iopub.status.idle": "2026-10-01T12:39:20.430612Z", "shell.execute_reply": "2026-10-01T12:39:20.430169Z" }, "lines_to_next_cell": 2 }, "outputs": [], "source": [ "path = simulate_path(lion, steps=5, seed=1234)" ] }, { "cell_type": "markdown", "id": "68056c2d", "metadata": { "lines_to_next_cell": 2 }, "source": [ "We could also provide a scheduling function to choose the actions ourselves. This is somewhat similar to the `bird` API." ] }, { "cell_type": "code", "execution_count": 3, "id": "be518c46", "metadata": { "execution": { "iopub.execute_input": "2026-10-01T12:39:20.432286Z", "iopub.status.busy": "2026-10-01T12:39:20.432103Z", "iopub.status.idle": "2026-10-01T12:39:20.435395Z", "shell.execute_reply": "2026-10-01T12:39:20.434747Z" }, "lines_to_next_cell": 0 }, "outputs": [], "source": [ "def scheduler(s: State) -> Action:\n", " return Action(\"rawr\")\n", "\n", "\n", "path2 = stormvogel.simulator.simulate_path(\n", " lion, steps=5, seed=1234, scheduler=scheduler\n", ")" ] }, { "cell_type": "markdown", "id": "15e37d92", "metadata": {}, "source": [ "We can also use the scheduler to create a partial model. This model contains all the states that have been discovered by the the simulation." ] }, { "cell_type": "code", "execution_count": 4, "id": "f24430f5", "metadata": { "execution": { "iopub.execute_input": "2026-10-01T12:39:20.437071Z", "iopub.status.busy": "2026-10-01T12:39:20.436891Z", "iopub.status.idle": "2026-10-01T12:39:20.462367Z", "shell.execute_reply": "2026-10-01T12:39:20.461732Z" } }, "outputs": [ { "data": { "text/html": [ "\n", "\n", "\n", " \n", " Network\n", " \n", " \n", " \n", " \n", " \n", "
\n", " \n", " \n", " \n", " \n", "\n" ], "text/plain": [ "" ] }, "metadata": {}, "output_type": "display_data" }, { "data": { "text/plain": [ "" ] }, "execution_count": 4, "metadata": {}, "output_type": "execute_result" } ], "source": [ "partial_model = stormvogel.simulator.simulate(\n", " lion, steps=5, scheduler=scheduler, seed=1234\n", ")\n", "show(partial_model)" ] }, { "cell_type": "markdown", "id": "fe22a51d", "metadata": {}, "source": [ "## Gymnasium-Compliant Environment\n", "\n", "Stormvogel models can be wrapped as a [Gymnasium](https://gymnasium.farama.org/)\n", "environment via `ModelEnv`. This lets you use standard reinforcement-learning\n", "libraries directly on a stormvogel MDP or DTMC without any manual glue code.\n", "\n", "`ModelEnv` supports both **MDP** and **DTMC** models:\n", "* For an MDP the action space is `Discrete(n_actions)`, one index per named action.\n", "* For a DTMC there is no choice, so the action space is `Discrete(1)` — always pass `0`.\n", "\n", "The observation space has two modes, selected by `obs_type`:\n", "* `\"index\"` (default): a plain integer — the index of the current state.\n", "* `\"valuations\"`: a `Dict` space built from variables that have a declared domain\n", " (`IntDomain`, `BoolDomain`, or `CategoricalDomain`), one `Discrete` component\n", " per variable." ] }, { "cell_type": "code", "execution_count": 5, "id": "288d20c4", "metadata": { "execution": { "iopub.execute_input": "2026-10-01T12:39:20.479514Z", "iopub.status.busy": "2026-10-01T12:39:20.479266Z", "iopub.status.idle": "2026-10-01T12:39:20.630066Z", "shell.execute_reply": "2026-10-01T12:39:20.629484Z" } }, "outputs": [ { "name": "stdout", "output_type": "stream", "text": [ "Observation space: Discrete(5)\n", "Action space: Discrete(3)\n", "Actions: [Action('hunt >:D'), Action('rawr'), Action(None)]\n" ] } ], "source": [ "from stormvogel.gym_env import ModelEnv\n", "\n", "env = ModelEnv(lion)\n", "print(\"Observation space:\", env.observation_space)\n", "print(\"Action space: \", env.action_space)\n", "print(\"Actions: \", env._index_to_action)" ] }, { "cell_type": "markdown", "id": "bcc171dd", "metadata": {}, "source": [ "The standard Gymnasium loop works as-is. `reset()` returns the initial\n", "observation and an info dict; `step(action)` returns the next observation,\n", "reward, terminated flag, truncated flag, and an info dict. The info dict\n", "always contains the raw `stormvogel.model.State` under the key `\"state\"`." ] }, { "cell_type": "code", "execution_count": 6, "id": "eb4653f3", "metadata": { "execution": { "iopub.execute_input": "2026-10-01T12:39:20.632116Z", "iopub.status.busy": "2026-10-01T12:39:20.631826Z", "iopub.status.idle": "2026-10-01T12:39:20.635869Z", "shell.execute_reply": "2026-10-01T12:39:20.635342Z" } }, "outputs": [ { "name": "stdout", "output_type": "stream", "text": [ "Initial state index: 0 — state: State(id=d0ecb70b-eb37-4419-bb5f-9e7ce877d141, labels=['init'])\n", "After 'hunt': obs = 0 | reward = 0.0 | terminated = False\n" ] } ], "source": [ "obs, info = env.reset(seed=42)\n", "print(\"Initial state index:\", obs, \"— state:\", info[\"state\"])\n", "\n", "hunt_idx = next(i for i, a in enumerate(env._index_to_action) if \"hunt\" in str(a))\n", "obs, reward, terminated, truncated, info = env.step(hunt_idx)\n", "print(\"After 'hunt': obs =\", obs, \"| reward =\", reward, \"| terminated =\", terminated)" ] }, { "cell_type": "markdown", "id": "8fd3612f", "metadata": {}, "source": [ "Passing an action that is not available in the current state raises\n", "`ActionUnavailableError` rather than silently producing incorrect behaviour.\n", "\n", "### Variable-domain observations\n", "\n", "When all variables in the model carry a declared domain, `obs_type=\"valuations\"`\n", "gives a `Dict` observation whose keys are variable names and whose values are\n", "non-negative integers (domain encoding). This is more informative than a raw\n", "state index and compatible with structured RL policies." ] }, { "cell_type": "code", "execution_count": 7, "id": "1e451239", "metadata": { "execution": { "iopub.execute_input": "2026-10-01T12:39:20.637594Z", "iopub.status.busy": "2026-10-01T12:39:20.637429Z", "iopub.status.idle": "2026-10-01T12:39:20.642341Z", "shell.execute_reply": "2026-10-01T12:39:20.641837Z" } }, "outputs": [ { "name": "stdout", "output_type": "stream", "text": [ "Observation space: Dict('car_pos': Discrete(4), 'chosen_pos': Discrete(4), 'reveal_pos': Discrete(4))\n", "Initial obs (all variables at sentinel -1): {'car_pos': 0, 'chosen_pos': 0, 'reveal_pos': 0}\n" ] } ], "source": [ "from stormvogel.examples.monty_hall import create_monty_hall_mdp\n", "\n", "mh = create_monty_hall_mdp()\n", "mh_env = ModelEnv(mh, obs_type=\"valuations\")\n", "print(\"Observation space:\", mh_env.observation_space)\n", "obs, info = mh_env.reset()\n", "print(\"Initial obs (all variables at sentinel -1):\", obs)" ] } ], "metadata": { "kernelspec": { "display_name": "Python 3 (ipykernel)", "language": "python", "name": "python3" }, "language_info": { "codemirror_mode": { "name": "ipython", "version": 3 }, "file_extension": ".py", "mimetype": "text/x-python", "name": "python", "nbconvert_exporter": "python", "pygments_lexer": "ipython3", "version": "3.13.15" }, "widgets": { "application/vnd.jupyter.widget-state+json": { "state": { "19b4c05dfa1a4e9e80efa0737d3be7c2": { "model_module": "@jupyter-widgets/base", "model_module_version": "2.0.0", "model_name": "LayoutModel", "state": { "_model_module": "@jupyter-widgets/base", "_model_module_version": "2.0.0", "_model_name": "LayoutModel", "_view_count": null, "_view_module": "@jupyter-widgets/base", "_view_module_version": "2.0.0", "_view_name": "LayoutView", "align_content": null, "align_items": null, "align_self": null, "border_bottom": null, "border_left": null, "border_right": null, "border_top": null, "bottom": null, "display": null, "flex": null, "flex_flow": null, "grid_area": null, "grid_auto_columns": null, "grid_auto_flow": null, "grid_auto_rows": null, "grid_column": null, "grid_gap": null, "grid_row": null, "grid_template_areas": null, "grid_template_columns": null, "grid_template_rows": null, "height": null, "justify_content": null, "justify_items": null, "left": null, "margin": null, "max_height": null, "max_width": null, "min_height": null, "min_width": null, "object_fit": null, "object_position": null, "order": null, "overflow": null, "padding": null, "right": null, "top": null, "visibility": null, "width": null } }, "6628725a6ecc48f18d7d111d05c7ca49": { "model_module": "@jupyter-widgets/output", "model_module_version": "1.0.0", "model_name": "OutputModel", "state": { "_dom_classes": [], "_model_module": "@jupyter-widgets/output", "_model_module_version": "1.0.0", "_model_name": "OutputModel", "_view_count": null, "_view_module": "@jupyter-widgets/output", "_view_module_version": "1.0.0", "_view_name": "OutputView", "layout": "IPY_MODEL_97ccecdc0d1e4bb6a35ae0b4f4fbe316", "msg_id": "", "outputs": [], "tabbable": null, "tooltip": null } }, "679c23e23ad34679a62940d7d1db862c": { "model_module": "@jupyter-widgets/base", "model_module_version": "2.0.0", "model_name": "LayoutModel", "state": { "_model_module": "@jupyter-widgets/base", "_model_module_version": "2.0.0", "_model_name": "LayoutModel", "_view_count": null, "_view_module": "@jupyter-widgets/base", "_view_module_version": "2.0.0", "_view_name": "LayoutView", "align_content": null, "align_items": null, "align_self": null, "border_bottom": null, "border_left": null, "border_right": null, "border_top": null, "bottom": null, "display": null, "flex": null, "flex_flow": null, "grid_area": null, "grid_auto_columns": null, "grid_auto_flow": null, "grid_auto_rows": null, "grid_column": null, "grid_gap": null, "grid_row": null, "grid_template_areas": null, "grid_template_columns": null, "grid_template_rows": null, "height": null, "justify_content": null, "justify_items": null, "left": null, "margin": null, "max_height": null, "max_width": null, "min_height": null, "min_width": null, "object_fit": null, "object_position": null, "order": null, "overflow": null, "padding": null, "right": null, "top": null, "visibility": null, "width": null } }, "6cef00808c024329aebcc166ae6a6f66": { "model_module": "@jupyter-widgets/output", "model_module_version": "1.0.0", "model_name": "OutputModel", "state": { "_dom_classes": [], "_model_module": "@jupyter-widgets/output", "_model_module_version": "1.0.0", "_model_name": "OutputModel", "_view_count": null, "_view_module": "@jupyter-widgets/output", "_view_module_version": "1.0.0", "_view_name": "OutputView", "layout": "IPY_MODEL_19b4c05dfa1a4e9e80efa0737d3be7c2", "msg_id": "", "outputs": [], "tabbable": null, "tooltip": null } }, "97ccecdc0d1e4bb6a35ae0b4f4fbe316": { "model_module": "@jupyter-widgets/base", "model_module_version": "2.0.0", "model_name": "LayoutModel", "state": { "_model_module": "@jupyter-widgets/base", "_model_module_version": "2.0.0", "_model_name": "LayoutModel", "_view_count": null, "_view_module": "@jupyter-widgets/base", "_view_module_version": "2.0.0", "_view_name": "LayoutView", "align_content": null, "align_items": null, "align_self": null, "border_bottom": null, "border_left": null, "border_right": null, "border_top": null, "bottom": null, "display": null, "flex": null, "flex_flow": null, "grid_area": null, "grid_auto_columns": null, "grid_auto_flow": null, "grid_auto_rows": null, "grid_column": null, "grid_gap": null, "grid_row": null, "grid_template_areas": null, "grid_template_columns": null, "grid_template_rows": null, "height": null, "justify_content": null, "justify_items": null, "left": null, "margin": null, "max_height": null, "max_width": null, "min_height": null, "min_width": null, "object_fit": null, "object_position": null, "order": null, "overflow": null, "padding": null, "right": null, "top": null, "visibility": null, "width": null } }, "a29100bd44c54f9089d50eb16d05e1be": { "model_module": "@jupyter-widgets/base", "model_module_version": "2.0.0", "model_name": "LayoutModel", "state": { "_model_module": "@jupyter-widgets/base", "_model_module_version": "2.0.0", "_model_name": "LayoutModel", "_view_count": null, "_view_module": "@jupyter-widgets/base", "_view_module_version": "2.0.0", "_view_name": "LayoutView", "align_content": null, "align_items": null, "align_self": null, "border_bottom": null, "border_left": null, "border_right": null, "border_top": null, "bottom": null, "display": null, "flex": null, "flex_flow": null, "grid_area": null, "grid_auto_columns": null, "grid_auto_flow": null, "grid_auto_rows": null, "grid_column": null, "grid_gap": null, "grid_row": null, "grid_template_areas": null, "grid_template_columns": null, "grid_template_rows": null, "height": null, "justify_content": null, "justify_items": null, "left": null, "margin": null, "max_height": null, "max_width": null, "min_height": null, "min_width": null, "object_fit": null, "object_position": null, "order": null, "overflow": null, "padding": null, "right": null, "top": null, "visibility": null, "width": null } }, "bcb48f7004184eb5b982aeb39f814e9a": { "model_module": "@jupyter-widgets/output", "model_module_version": "1.0.0", "model_name": "OutputModel", "state": { "_dom_classes": [], "_model_module": "@jupyter-widgets/output", "_model_module_version": "1.0.0", "_model_name": "OutputModel", "_view_count": null, "_view_module": "@jupyter-widgets/output", "_view_module_version": "1.0.0", "_view_name": "OutputView", "layout": "IPY_MODEL_a29100bd44c54f9089d50eb16d05e1be", "msg_id": "", "outputs": [], "tabbable": null, "tooltip": null } }, "f6a6c9345b724f88838454f0933d651d": { "model_module": "@jupyter-widgets/output", "model_module_version": "1.0.0", "model_name": "OutputModel", "state": { "_dom_classes": [], "_model_module": "@jupyter-widgets/output", "_model_module_version": "1.0.0", "_model_name": "OutputModel", "_view_count": null, "_view_module": "@jupyter-widgets/output", "_view_module_version": "1.0.0", "_view_name": "OutputView", "layout": "IPY_MODEL_679c23e23ad34679a62940d7d1db862c", "msg_id": "", "outputs": [], "tabbable": null, "tooltip": null } } }, "version_major": 2, "version_minor": 0 } } }, "nbformat": 4, "nbformat_minor": 5 }