Ikechukwu Daniel Adebi, Peter Stone, Mitchell Pryor
The University of Texas at Austin
- Create the conda environment from
environment_lcamp.yml:conda create -f environment_lcamp.yml conda activate lcamp
- Build the LCAMP dataset (only needed once per dataset version), see
process_lcamp_dataset.py. By default it writes to./lcamp_dataset(override with its own--data_dir), producing a directory containingtrain_dataset.json,test_dataset.json, andintrinsics.npy, and that directory is what's passed below as--data_root/--data_dirintrain.py/eval.py. - Get the background images:
LCAMPDataset(inlcamp_dataset.py) unconditionally loads a background image list on init (used for the--randomize_background/--random_backgroundsaugmentation, but read regardless), sourced from the MIT Indoor 67 scene dataset. From that page, download the images archive and theTrainImages.txt/TestImages.txtsplit files, then lay them out as:relative to wherever you runbackgrounds/mit_indoor_67/TrainImages.txt backgrounds/mit_indoor_67/TestImages.txt backgrounds/mit_indoor_67/raw/Images/<category>/<image>.jpgtrain.py/eval.pyfrom (or under$SCRATCH/backgrounds/mit_indoor_67/...if using--cluster_loc stampede3, see the--cluster_locnote below).
All commands below (train.py, eval.py) assume the lcamp environment is
active.
A pretrained checkpoint is available on HuggingFace:
danieladebi/lcamp-model.
Download it into checkpoints/ to skip training and go straight to
evaluation.
Training is done with train.py via torchrun (DDP). The dataset must already
be built (see process_lcamp_dataset.py) as a directory containing
train_dataset.json, test_dataset.json, and intrinsics.npy.
Set --nproc_per_node to the number of GPUs on your machine (use 1 for a
single-GPU/local run):
torchrun --nproc_per_node=<NUM_GPUS> train.py \
--save_path checkpoints/my_lcamp_model \
--data_root ./lcamp_dataset \
--model_type resnet \
--epochs 50 \
--batch_size 64 \
--lr 1e-3 \
--min_lr 1e-5 \
--use_camera_frame \
--use_depth \
--use_text_instructions \
--emphasize_part_mask \
--bbox_loc \
--cluster_loc <CLUSTER_LOC> \
--lambda_axis 8 \
--lambda_anchor 8 \
--lambda_joint 1Notable flags (train.py --help for the complete list):
--model_type:resnet(default LCAMP backbone) ordino.--use_text_instructions: enables the language-conditioned model (LCAMPModelResnetLanguage); required if you plan to run thelcampbackend with instructions inlcamp_service_node.py.--use_depth: adds depth as an input channel.--bbox_loc: predict axis location as(u,v,z)decoded via the part bounding box instead of raw XYZ, must match--bbox_locineval.pyanduse_bbox_locinlcamp_service_node.py/lcamp.launch.pyfor the same checkpoint. Whenever this is active, pass--bbox_locexplicitly so the flag stays consistent acrosstrain.py/eval.py/deployment.--resume/--resume_from <path>: resume fromcheckpoint_best.ptin--save_path, or from an explicit checkpoint path.--save_freq,--eval_freq: checkpoint/eval cadence in epochs.--cluster_loc: the dataset JSON (fromprocess_lcamp_dataset.py) stores absolute file paths for the rgb/depth/mask files, so if you're loading the dataset on a different machine/path layout than where it was built, those paths won't resolve.nrg(default) uses the JSON paths as-is,stampede3is a legacy alias for$SCRATCH. For anything else, just pass your own dataset root directory as--cluster_loc /path/to/your/root, andlcamp_ dataset.pywill swap it in for the legacy/storage/danieladebiprefix in the rgb/depth/mask paths, and root the background images (see above) under it too. No code changes needed.
Checkpoints are written to --save_path as checkpoint_best.pt,
checkpoint_last.pt, and periodic checkpoint_{epoch}.pt files.
Evaluation is done with eval.py (single process, no DDP):
python eval.py \
--model_path checkpoints/my_lcamp_model/checkpoint_best.pt \
--data_dir ./lcamp_dataset \
--split test \
--model_type resnet \
--use_camera_frame \
--use_depth \
--use_text_instructions \
--bbox_loc \
--cluster_loc <CLUSTER_LOC> \
--visualizeThis must use the same architecture flags used at training time
(--model_type, --use_depth, --use_text_instructions,
--exclude_object_mask, --emphasize_part_mask/--emphasize_object_mask,
--bbox_loc, etc.), mismatches will load state dict weights incorrectly or
error, and --use_camera_frame/--cluster_loc must match how the dataset
was built.
Notable flags:
--split:trainortest.--visualize: saves per-sample prediction images.--random_preds: baseline using random axis/location predictions.--free_motion_query_prob: fraction of eval samples using a free-motion (zero-vector) query instead of the GT axis, to measure reliance on visual cues vs. the GT axis itself.
Results (per-object/per-category success rates, axis/location error, joint
type accuracy) are written to results/<model_name>/, where <model_name>
is derived from --model_path.
Once you have a checkpoint you're happy with, drop it into checkpoints/
and point the model_name launch argument in launch/lcamp.launch.py at it
(see the launch file for details on use_bbox_loc and other runtime flags).
Running predictions on a real Boston Dynamics Spot (via the ROS2 service node
in scripts/lcamp_service_node.py) is outside the scope of this README, see
ros2_setup/ROS2_SETUP.md for that setup.
TODO: fill in once paper is completeThis project is licensed under the Apache License 2.0. See LICENSE for details.