RoboCasa365#

RoboCasa365 environment overview

Kitchen scenes, objects, and tasks in RoboCasa365. Source: RoboCasa365 project.#

Run kitchen manipulation tasks with RPent in RoboCasa365, then reproduce Target50 experiments. This integration uses the PandaOmron mobile manipulator and the RLDX-1 action model; its CLI name is robocasa.

Overview#

Check the model, task, and runtime requirements before following the installation and run steps.

Action Models

RLDX-1

Planners

api, claude_code, codex

Tasks

Target50 kitchen tasks

Hardware

Linux, NVIDIA GPU; Python 3.10; CUDA and EGL.

Tasks#

Target50 covers the following categories. The reproduction section lists the complete task/seed matrix and run limits.

Category

Tasks

Scope

Atomic

18

Individual kitchen operations.

Composite-Seen

16

Composite tasks in seen categories.

Composite-Unseen

16

Composite tasks in unseen categories.

Observation and Action#

The table distinguishes planner tools, model inputs, and the environment’s success criterion.

Item

Description

Observation

Camera views, depth/world coordinates, and mobile-base, end-effector, and gripper state. RLDX-1 receives three camera histories plus state and task text.

Action

The planner calls rldx_skill and motion primitives. RLDX-1 predicts end-effector, gripper, base-motion, and control-mode commands.

Reward / success

Evaluate success using the environment’s _check_success() result exposed as state.success.

Task prompt

Use the current environment’s complete task_language; historical memory does not replace it.

Installation and Resources#

Use Linux, an NVIDIA GPU, a working CUDA/EGL setup, git, and uv. If you already have the repository, enter it and start at environment creation.

RLDX-1 requires Python 3.10. Create a dedicated environment and install the complete RoboCasa365 stack with .[robocasa]:

git clone https://github.com/RLinf/RPent.git
cd RPent
uv venv --python 3.10 .venv-robocasa
source .venv-robocasa/bin/activate

First install a matching CUDA-enabled PyTorch and torchvision pair using the PyTorch installation selector for your GPU, driver and Python version. Run the selected command in this environment (use uv pip in place of pip). The RLDX dependency requires Torch >= 2.7 and torchvision >= 0.22; choose a mutually compatible pair, not two independent versions. Then install RPent:

uv pip install -e ".[robocasa]" \
   --constraint robots/robocasa/eval/target50-constraints.txt
uv pip check

The constraints file pins compatibility-sensitive Target50 dependencies. The robocasa extra installs the rpent branches of RoboCasa, RLDX, and Robosuite. Do not also install rlinf-robocasa365, which provides the same import package.

Choose Torch, torchvision, and CUDA for your machine. The reference run used Torch 2.7.0, torchvision 0.22.0, and CUDA 12.6. Record resolved dependencies and Git revisions for every reproduction because branches can advance:

uv pip freeze > installed-requirements.txt
git rev-parse HEAD > rpent-revision.txt

flash-attn is optional; RLDX-1 uses PyTorch SDPA when it is absent. If needed, follow the FlashAttention instructions to select a compatible build.

Post-install setup

Download the kitchen assets (~10 GB) outside site-packages so they survive reinstalls. Target50 does not use RoboCasa dataset or teleop macros, so skip the optional private-macros setup:

robocasa-download-assets --assets-path ~/.robocasa/assets --no-macros -y

It prints the environment variable to export afterwards; add it to the shell that launches rpent:

export ROBOCASA_ASSETS_PATH=~/.robocasa/assets

The installer supplies downloaded collections and bundled scene files. Add --skip-existing on subsequent runs to check existing downloads. See troubleshooting below for resource conflicts, disk space, and camera errors.

RLDX-1 checkpoint

The --vla-model-path flag on the run commands below expects a local path to the RLDX-1-FT-RC365 checkpoint (the RoboCasa365 fine-tune). Download it from HuggingFace:

hf download RLWRLD/RLDX-1-FT-RC365 \
   --revision 587e9ecdcc5e7184fcc17f58713908edff5af041 \
   --local-dir ./checkpoints/rldx-1-ft-rc365

If the download is slow, use the HF mirror:

HF_ENDPOINT=https://hf-mirror.com hf download RLWRLD/RLDX-1-FT-RC365 \
   --revision 587e9ecdcc5e7184fcc17f58713908edff5af041 \
   --local-dir ./checkpoints/rldx-1-ft-rc365

RLDX-1 backbone support files

The FT checkpoint contains the weights but also references RLWRLD/RLDX-1-VLM for architecture, processor and tokenizer metadata. Target50 freezes revision 4b9f870d1287e0d38d7eb1445e6d8c60afe66dd7: 15 non-weight files, about 16.4 MB including documentation and images. Download these into the same cache used when launching RPent:

export HF_HOME="$PWD/.cache/huggingface"
export HF_HUB_CACHE="$HF_HOME/hub"
hf download RLWRLD/RLDX-1-VLM \
   --revision 4b9f870d1287e0d38d7eb1445e6d8c60afe66dd7 \
   --include "*.json" "*.txt" "*.jinja" "*.md" "*.png" ".gitattributes" \
   --exclude "*.safetensors.index.json"

No additional base weights are required. Keep these cache variables in the launch shell; do not shadow them with an empty TRANSFORMERS_CACHE. The RoboCasa VLA worker automatically uses this same support revision for both ordinary and Target50 runs, including separately started RPent VLA servers. There is no extra revision flag or manual cache-ref edit. The pin applies only to backbone metadata, not the weights selected by --vla-model-path. Model and asset licenses apply separately from RPent’s code license.

Run a Task#

Configure your model service with Planner Configuration and check it using rpent-check-llm. Run OpenDrawer with seed 1:

rpent --robot robocasa \
      --task-name OpenDrawer \
      --split target \
      --seed 1 \
      --vla-model-path ./checkpoints/rldx-1-ft-rc365 \
      --planner claude_code \
      --model claude-opus-4-8

RoboCasa does not select a planner implementation; the api, claude_code, and codex planners can run this robot. See Planner Configuration for configuration.

View Results#

Task success comes from the environment’s _check_success() result, exposed as state.success. The planner’s finish status ends its conversation and is not an evaluation label. Inspect result.json, transcript_*.json, and run.log in the output directory; service startup errors are in env_server.log and vla_server.log.

Use --dashboard to watch cameras and planner output; see Interactive Usage for the shared workflow.

Task Memory#

With --memory-profile hf (the default), the CLI and Dashboard synchronize robocasa/** from the current main branch of the RLinf/RPent-memory dataset. Memory is not pinned to a commit. The layout is:

memory/robocasa/
├── task-specific/
│   ├── <Task>_s0.json
│   ├── <Task>_s0_recipe.jsonl
│   └── <Task>.md              # optional
└── global/
    └── GLOBAL_MEMORY.md

In HF evaluation, RoboCasa provides the current task’s available JSON, recipe and Markdown alongside global/GLOBAL_MEMORY.md. The planner uses read_text_file to consult relevant task-specific and global guidance as needed. It chooses when and how much to read; actions and completion do not require every file to be read first. There is no option to disable the global layer.

The prompt and file tools share the same selection. RPent file tools deny other tasks’ memory; this is a tool restriction, not an operating-system sandbox. A missing JSON/JSONL pair is allowed: the planner continues with live observations and global guidance. A half-present pair is an error. Missing optional Markdown is logged, and the global file must exist. Files are discovered by task name and directory; no extra index is needed. The CLI validates memory before starting robot services. Dashboard validates the memory root and global layer before starting its shared VLA, then checks each selected task’s files before starting that task’s environment. A task memory error leaves the existing shared VLA available for other tasks.

Live task_language, RGB-D observations, task progress and tool results take precedence over memory. Apply a global strategy only when its visible preconditions hold. Continue VLA calls while contact, a held object, fixture progress or a counter increase shows progress. After two consecutive calls without contact or visible progress, re-ground and make a bounded pose adjustment. Every VLA call uses the full, verbatim live task language. Historical vla_act entries describe strategies; use current tools and never replay historical coordinates. Reset remains unavailable during evaluation.

To use local memory, download into a fresh directory and select the local profile. This also avoids retaining deleted files in an older download directory:

hf download RLinf/RPent-memory --repo-type dataset \
   --include 'robocasa/**' --local-dir ./target50-memory

rpent --robot robocasa \
      --task-name OpenDrawer --split target --seed 1 \
      --vla-model-path /path/to/rldx \
      --planner codex --model gpt-5.5 --reasoning-effort xhigh \
      --memory-profile local --memory-dir ./target50-memory/robocasa

Future memory updates are published to HF main. The reproduce/memory archive retains the GPT-5.5 Harness-VLA reproduction resources at d8c25a7f, with its existing contents and directory names unchanged. Its RoboCasa files use task_only/; the current loader requires task-specific/ and does not convert the old layout. Downloading that archive and passing it to the current --memory-profile local loader is not a supported reproduction command. The archive’s README describes its historical behavior, not the current CLI. No matching RoboCasa code/data snapshot is established by this guide; the commands above use the current main corpus.

Local exploration output is also supported directly, without conversion. For --task-name <Task> --split <split>, local evaluation selects either the published <Task>_s0.json / <Task>_s0_recipe.jsonl pair or the native <Task>_<split>_s0.json / <Task>_<split>_s0_recipe.jsonl pair in task-specific/. Either half-pair is an error. If both pairs exist, use separate --memory-dir directories; RPent does not choose between them. When neither pair exists, evaluation can use global guidance alone.

Local evaluation also exposes global/*.md and the task-family/*.md leaves whose YAML frontmatter matches suite: robocasa, regime: <split> and task_id: <Task>. Other tasks and splits are excluded. An optional task-specific/<Task>.md remains available. Evaluation neither requires nor exposes the corpus-wide MEMORY.md index or _internal/; it lists the selected files directly. Missing global memory prevents evaluation startup, including with the local profile. The same selection drives prompts, file permissions, and read audits; results record the actual profile and matching family identity for validation.

Custom Memory Sources#

To use another HF dataset with the same robocasa/ layout, set RPENT_MEMORY_HF_REPO=<owner>/<dataset> when launching RPent with --memory-profile hf. This accepts a dataset repository ID, not a browser URL.

For a custom subdirectory or a maintained branch, download that subtree into a fresh directory and use the local profile:

hf download <owner>/<dataset> --repo-type dataset \
   --include '<subpath>/**' --local-dir ./custom-memory

# Add to the RPent run command:
# --memory-profile local --memory-dir ./custom-memory/<subpath>

The selected directory uses task-specific/ and must provide at least one readable global/*.md file. Add --revision <branch> to the HF download command when selecting a branch. No delivery package or migration script is required.

Exploration Mode#

Add --explore to let the planner retry a task across fresh episodes and write local memory. The directory may start empty; exploration can read its index and write to its own inbox. As with LIBERO, one run allows up to three planner sessions with at most five attempts per session by default:

rpent --robot robocasa --task-name OpenDrawer --split target --seed 0 \
  --vla-model-path /path/to/rldx \
  --planner codex --reasoning-effort high --planner-timeout-s 7200 \
  --explore --explore-sessions 3 --explore-attempts-per-session 5 \
  --memory-dir /path/to/robocasa-memory

reset uses the environment’s ordinary episode reset. The runner exports only the winning commands after the final reset. Exploration memory is written to the current local inbox. Drafts are merged when the run completes normally without an agent execution error. Pass --no-auto-merge-memory to disable automatic merging. The exploration prompt is in robots/robocasa/prompts/explore.py and covers mobile-base use, task_progress, RLDX continuity, and failed-attempt notes.

Shared memory merge publishes accepted proposals under task-family/ and global/, copies the successful audit/recipe pair into task-specific/, and refreshes MEMORY.md. To evaluate these native files, use seed-0 exploration output and --memory-profile local with the same directory. Evaluation requires global memory to have been published first and uses its own task/split access boundary; exploration keeps its retry and inbox workflow.

Experiment Reproduction (Target50)#

The current robots/robocasa/eval/target50_v2.json protocol (robocasa-harness-vla-v2) uses task/global memory without pinning its data version. It preserves the target task/seed matrix, cell time limits, no-reset rule, environment success predicate, and 40/999/8 RLDX settings. The protocol ID identifies the result format and evaluation rules; it lets the validator distinguish v1 from v2 and does not select a memory data version.

Results record the fixed task/global selection, missing files and actual reads. The validator accepts zero or partial reads, while checking task boundaries and the audit structure. Missing or corrupt audit files are reported separately; read completeness does not determine the environment result’s validity or success. Each run starts a fresh audit, even when reusing an output directory.

Memory contents are not compared across runs. HF main receives future updates, while reproduce/memory remains an unchanged historical archive. For a repeatable comparison using the current layout, download memory once and use the same unchanged directory with --memory-profile local --memory-dir for every cell. Retain the files and record the HF commit or hashes in local experiment notes. RPent does not pin memory or add data revision identifiers to result metadata.

The manifests describe the evaluation matrix and validation rules; the leaderboard displays independently reported scores. A 340-cell result alone does not establish which memory, model or code configuration produced it. The current v2 manifest includes a GPT-5.5 reference profile; it is not a universal validator for every model on the leaderboard.

  • target50.json retains the historical v1 task-specific protocol. Validate compatible historical records with --manifest robots/robocasa/eval/target50.json.

  • target50_v2.json describes current runs with task-specific and global memory. It is the default for new results and validation.

  • Published leaderboard scores retain their original reported sources; they are not reclassified as v2 results without matching run evidence.

Each cell is one task/seed combination. Source dependencies follow the recorded rpent branches; retain their resolved commits with the experiment artifacts.

RoboCasa Target50 matrix#

Split

Tasks

Seed range per task

Cell timeout

Cells

Atomic

18

1–10

1800 s

180

Composite-Seen

16

1–5

3600 s

80

Composite-Unseen

16

1–5

3600 s

80

Total

50

340

Seen/unseen describes whether tasks occur in the pretraining data; target kitchens form a separate held-out scene split. See the RoboCasa dataset definitions. The tasks split into three groups:

  • Atomic (18) — single-primitive articulation and pick-place tasks: CloseBlenderLid, CloseFridge, CloseToasterOvenDoor, CoffeeSetupMug, NavigateKitchen, OpenCabinet, OpenDrawer, OpenStandMixerHead, PickPlaceCounterToCabinet, PickPlaceCounterToStove, PickPlaceDrawerToCounter, PickPlaceSinkToCounter, PickPlaceToasterToCounter, SlideDishwasherRack, TurnOffStove, TurnOnElectricKettle, TurnOnMicrowave, TurnOnSinkFaucet.

  • Composite seen (16) — multi-step tasks represented in the pretraining data: ScrubCuttingBoard, StackBowlsCabinet, WashLettuce, RinseSinkBasin, PreSoakPan, StirVegetables, LoadDishwasher, SteamInMicrowave, SetUpCuttingStation, GetToastedBread, DeliverStraw, KettleBoiling, PrepareCoffee, StoreLeftoversInBowl, SearingMeat, PackIdenticalLunches.

  • Composite unseen (16) — multi-step tasks absent from the pretraining data: ArrangeBreadBasket, ArrangeTea, BreadSelection, CategorizeCondiments, CuttingToolSelection, GarnishPancake, GatherTableware, HeatKebabSandwich, MakeIceLemonade, PanTransfer, PortionHotDogs, RecycleBottlesByType, SeparateFreezerRack, WaffleReheat, WashFruitColander, WeighIngredients.

Pass any of these to --task-name. The full RoboCasa catalog is larger; see the RoboCasa upstream.

For a repeatable Target50 comparison, prepare the current corpus once using the Task Memory download command above. Keep that local directory unchanged for all cells and retain its source revision or hashes with the results.

For Target50, first prepare the resources and local memory above, then invoke one ordinary command for each manifest cell. The Codex reference profile is gpt-5.5, xhigh, and max_turns=100; RoboCasa itself remains planner-agnostic. For the scene identity, use the ordinary --seed argument and do not set RLDX_RESET_SEED. Ordinary RoboCasa uses max_chunks=70; Target50 alone overrides it to 40. Freeze the Target50 RLDX execution values first:

export RLDX_MAX_CHUNKS=40
export RLDX_SETTLE_PATIENCE=999
export RLDX_ACTION_STEPS_PER_CHUNK=8
unset RLDX_RESET_SEED

The first OpenDrawer Atomic cell is:

rpent --robot robocasa \
      --task-name OpenDrawer --split target --seed 1 \
      --vla-model-path ./checkpoints/rldx-1-ft-rc365 --cuda-device 0 \
      --planner codex --model gpt-5.5 --reasoning-effort xhigh \
      --max-turns 100 --planner-timeout-s 1800 \
      --memory-profile local \
      --memory-dir ./target50-memory/robocasa \
      --output-dir ./runs/target50/atomic/OpenDrawer_s1

Use --planner-timeout-s 3600 for either composite split. Execute Atomic, Composite-Seen, and Composite-Unseen in that order. A cell succeeds only when the final recorded environment state has state.success=true; the planner’s finish(status=...) argument is not an evaluation label. Valid task failures and planner timeouts are not retried. Retry an infrastructure failure only when no valid environment result was produced for that cell.

Every completed command atomically writes <output-dir>/result.json using the final environment state.success. The record includes the effective protocol values but omits provider errors and credentials. Once all cells are present under <results-root>/<manifest-split>/<Task>_s<seed>/result.json, validate the fixed denominator and print the task-weighted score with:

python -m robots.robocasa.eval.validate_target50 ./runs/target50

Reported and Historical Target50 Results#

The leaderboard is the source for the reported rates below. The RPent configurations are Codex / GPT-5.5 / xhigh / reasoning, Codex / GPT-6 Astra / low / reasoning, and Claude Code / Opus-4.7 / max.reasoning. The Harness VLA reference column reports GPT-5.5 results from paper Table 4. Overall weights all 50 tasks equally; it is not the fraction of successful cells among 340.

Reported Target50 success rates#

Split

RPent / GPT-5.5

RPent / GPT-6 Astra

RPent / Opus-4.7

Harness VLA / GPT-5.5 reference

Atomic-Seen

92.0%

87.78%

79.4%

92.0%

Composite-Seen

61.0%

43.75%

47.5%

61.0%

Composite-Unseen

13.8%

42.50%

15.0%

13.8%

Overall (task-weighted)

57.1%

59.20%

48.6%

57.1%

The Astra entry reports 59.20% Overall, with 87.78% / 43.75% / 42.50% for the three splits. Its episode count is 340 (180/80/80) following the contributor-confirmed correction. The earlier 250-cell information was an unsynchronized historical record. The correction preserves reported rates; it does not infer success counts from rounded rates or claim a new audit of all 340 original results.

Historical Codex Reproduction#

The archived reproduction contains all 340 cells and reports the following task-level aggregates. These historical values do not describe a new v2 run:

Codex Target50 reproduction#

Split

Successful cells

Success rate

Harness VLA reference

Atomic

163/180

90.56%

165/180 (91.67%)

Composite-Seen

49/80

61.25%

45/80 (56.25%)

Composite-Unseen

12/80

15.00%

11/80 (13.75%)

Overall (task-weighted)

N/A

57.00%

55.40%

The archived per-task table contains the success count and accuracy for every task. This historical record is task-level aggregate data; it does not include per-seed traces, raw trajectories, or failure classifications and therefore is not a per-cell audit artifact.

Environment Checks#

After installing RoboCasa and its assets, run the opt-in environment smoke suite to check simulator installation and interfaces. It requires no planner credentials or VLA checkpoint:

uv pip install pytest pytest-timeout
RPENT_RUN_ROBOCASA_INTEGRATION=1 \
   pytest tests/integration_tests/robots/robocasa/test_target50_runtime_smoke.py -v

The four cases cover OpenDrawer, NavigateKitchen, and PickPlaceCounterToCabinet at seed 1, plus mobile-camera movement. Task checks verify construction/reset, 12D actions, operation cameras, navigation RGB-D/world map, the success predicate, and clean close. The camera check verifies pose and image changes after eight base steps. These real-simulator tests require a working GPU/EGL setup and are separate from offline CPU CI; skipped tests are not passes.

Troubleshooting#

The asset root must contain downloaded collections and bundled scene, arena, and fixture files. Preserve attribution files. --skip-existing checks download inventories; for conflicting files, confirm that replacement is intended before using:

robocasa-download-assets --assets-path ~/.robocasa/assets --no-macros --overwrite -y

--overwrite takes precedence over --skip-existing and replaces only files in the installation scope. By default, conflicting files are preserved and the destination must support hard links. Allow space for ZIP archives and unpacked data, plus existing data during replacement. Rerun the download after an interruption.

Run the environment smoke tests first. After downloading all four resources, use the existing RoboCasa E2E component test to verify VLA worker startup, HTTP RPC and first inference. Install .[test] if needed, select one available GPU and use a fresh output directory:

CUDA_VISIBLE_DEVICES=0 \
RLDX_MODEL_PATH="$PWD/checkpoints/rldx-1-ft-rc365" \
RPENT_E2E_OUTPUT_DIR="$PWD/e2e-robocasa" \
HF_HUB_OFFLINE=1 TRANSFORMERS_OFFLINE=1 NO_ALBUMENTATIONS_UPDATE=1 \
MUJOCO_GL=egl python -m pytest -q \
   tests/e2e_tests/robocasa/test_components.py::test_rldx_component --timeout=300

These checks do not run a planner or create benchmark results. Skipped tests are not passes. Offline variables apply only to the check; ordinary HF memory sync needs network access. Keep remote planner proxies unchanged. The lightweight protocol tests still validate all 50 tasks and the fixed 340-cell denominator; full benchmark execution is a separate procedure.

  • For slow package downloads, use UV_HTTP_TIMEOUT=600 and put caches and temporary files on a sufficiently large filesystem. Retry the pinned HF download; apparent shard size is not a completeness check. Do not disable TLS.

  • A read-only asset failure needs the corrected RoboCasa dependency, not writable canonical assets. Transformed XML uses the temporary directory.

  • For an RLDX offline cache miss, check the support snapshot and cache variables above. NO_ALBUMENTATIONS_UPDATE=1 disables only an import-time version check, not image processing. Keep the existing image-geometry fallback.

  • Test the selected Torch/CUDA build with a GPU operation and EGL render, not just the driver’s version display. Use a build compatible with the host.

  • For shared read-only installations, set NUMBA_CACHE_DIR to a writable per-user directory instead of making package code writable.

  • If navigation RGB-D or world-map rendering reports a missing mobilebase0_navview, reinstall .[robocasa] to refresh the RLinf/robosuite rpent branch. Do not patch installed XML files manually.

  • If read_text_file reports a missing current-task result, check the memory/robocasa/task-specific/ corpus or the selected local directory. RPent does not fall back to another task’s memory. Markdown is optional; Atomic tasks have no published <Task>.md.

  • Environment and VLA startup failures are recorded in <output_dir>/env_server.log and <output_dir>/vla_server.log; also inspect <output_dir>/run.log for the run-level error.

  • Only the exact 127.0.0.1 and localhost hostnames bypass HTTP proxies automatically. Other hostnames and IPs use the standard proxy environment; add the exact host to NO_PROXY and no_proxy only when it should be reached directly.

Implementation Notes#

The RoboCasa toolkit exposes the same shape of tools as LIBERO (a primitive call, a state view, a finish), with two RoboCasa-specific aspects:

  • Env-side helpers. Grasp checks and action assembly need the live simulator env, so they live in env_server as RPCs. The agent-side skill holds both clients: the env client for render/step, the model client for RLDX-1 inference. See Add a Robot or Simulator for the rationale.

  • Observation shape. RLDX-1 sees 3 camera video tensors (1, T, H, W, 3) stacked over history T, plus state.* and annotation.* fields. The session id is not part of the observation — it is managed automatically by the RPC framework: RpcClient generates a private rpc_ + uuid hex session id, wait_for_ready registers it with the server on connect; the server tracks each session’s idle time and a background sweep thread reaps sessions idle longer than the timeout (default 3600s), and the client sends session.close via atexit on process exit. Business code (rldx_skill / vla_client) never sees the session id directly; the server injects it into predict / reset_session to isolate per-client RLDX memory/RTC policy state.