Upload config

Browse files

Files changed (3) hide show

README.md +199 -0
config.json +222 -0
prismatic_config.py +309 -0

README.md ADDED Viewed

	@@ -0,0 +1,199 @@

+---
+library_name: transformers
+tags: []
+---
+# Model Card for Model ID
+<!-- Provide a quick summary of what the model is/does. -->
+## Model Details
+### Model Description
+<!-- Provide a longer summary of what this model is. -->
+This is the model card of a 🤗 transformers model that has been pushed on the Hub. This model card has been automatically generated.
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+### Model Sources [optional]
+<!-- Provide the basic links for the model. -->
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+## Uses
+<!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
+### Direct Use
+<!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
+[More Information Needed]
+### Downstream Use [optional]
+<!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
+[More Information Needed]
+### Out-of-Scope Use
+<!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
+[More Information Needed]
+## Bias, Risks, and Limitations
+<!-- This section is meant to convey both technical and sociotechnical limitations. -->
+[More Information Needed]
+### Recommendations
+<!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+## How to Get Started with the Model
+Use the code below to get started with the model.
+[More Information Needed]
+## Training Details
+### Training Data
+<!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
+[More Information Needed]
+### Training Procedure
+<!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
+#### Preprocessing [optional]
+[More Information Needed]
+#### Training Hyperparameters
+- **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
+#### Speeds, Sizes, Times [optional]
+<!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
+[More Information Needed]
+## Evaluation
+<!-- This section describes the evaluation protocols and provides the results. -->
+### Testing Data, Factors & Metrics
+#### Testing Data
+<!-- This should link to a Dataset Card if possible. -->
+[More Information Needed]
+#### Factors
+<!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
+[More Information Needed]
+#### Metrics
+<!-- These are the evaluation metrics being used, ideally with a description of why. -->
+[More Information Needed]
+### Results
+[More Information Needed]
+#### Summary
+## Model Examination [optional]
+<!-- Relevant interpretability work for the model goes here -->
+[More Information Needed]
+## Environmental Impact
+<!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+## Technical Specifications [optional]
+### Model Architecture and Objective
+[More Information Needed]
+### Compute Infrastructure
+[More Information Needed]
+#### Hardware
+[More Information Needed]
+#### Software
+[More Information Needed]
+## Citation [optional]
+<!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
+**BibTeX:**
+[More Information Needed]
+**APA:**
+[More Information Needed]
+## Glossary [optional]
+<!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
+[More Information Needed]
+## More Information [optional]
+[More Information Needed]
+## Model Card Authors [optional]
+[More Information Needed]
+## Model Card Contact
+[More Information Needed]

config.json ADDED Viewed

	@@ -0,0 +1,222 @@

+{
+  "auto_map": {
+    "AutoConfig": "prismatic_config.TrajectoryVLAConfig"
+  },
+  "cheat": false,
+  "model_type": "trajectoryvla",
+  "num_timesteps": 6,
+  "prismatic_config": {
+    "_name_or_path": "",
+    "add_cross_attention": false,
+    "arch_specifier": "no-align+gelu-mlp",
+    "architectures": [
+      "TrajectoryVLA"
+    ],
+    "auto_map": {
+      "AutoModelForVision2Seq": "prismatic_model.TrajectoryVLA"
+    },
+    "bad_words_ids": null,
+    "begin_suppress_tokens": null,
+    "bos_token_id": null,
+    "chunk_size_feed_forward": 0,
+    "cross_attention_hidden_size": null,
+    "decoder_start_token_id": null,
+    "diversity_penalty": 0.0,
+    "do_sample": false,
+    "early_stopping": false,
+    "encoder_no_repeat_ngram_size": 0,
+    "eos_token_id": null,
+    "exponential_decay_length_penalty": null,
+    "finetuning_task": null,
+    "forced_bos_token_id": null,
+    "forced_eos_token_id": null,
+    "hf_llm_id": "meta-llama/Llama-2-7b-hf",
+    "id2label": {
+      "0": "LABEL_0",
+      "1": "LABEL_1"
+    },
+    "image_resize_strategy": "letterbox",
+    "image_sizes": [
+      224,
+      224
+    ],
+    "is_decoder": false,
+    "is_encoder_decoder": false,
+    "label2id": {
+      "LABEL_0": 0,
+      "LABEL_1": 1
+    },
+    "length_penalty": 1.0,
+    "llm_backbone_id": "llama2-7b-pure",
+    "llm_max_length": 2048,
+    "max_length": 20,
+    "min_length": 0,
+    "model_type": "prismatic",
+    "no_repeat_ngram_size": 0,
+    "num_beam_groups": 1,
+    "num_beams": 1,
+    "num_return_sequences": 1,
+    "output_attentions": false,
+    "output_hidden_states": false,
+    "output_projector_states": false,
+    "output_scores": false,
+    "pad_to_multiple_of": 64,
+    "pad_token_id": 32000,
+    "prefix": null,
+    "problem_type": null,
+    "pruned_heads": {},
+    "remove_invalid_values": false,
+    "repetition_penalty": 1.0,
+    "return_dict": false,
+    "return_dict_in_generate": false,
+    "sep_token_id": null,
+    "suppress_tokens": null,
+    "task_specific_params": null,
+    "temperature": 1.0,
+    "text_config": {
+      "_name_or_path": "",
+      "add_cross_attention": false,
+      "architectures": null,
+      "attention_bias": false,
+      "attention_dropout": 0.0,
+      "bad_words_ids": null,
+      "begin_suppress_tokens": null,
+      "bos_token_id": 1,
+      "chunk_size_feed_forward": 0,
+      "cross_attention_hidden_size": null,
+      "decoder_start_token_id": null,
+      "diversity_penalty": 0.0,
+      "do_sample": false,
+      "early_stopping": false,
+      "encoder_no_repeat_ngram_size": 0,
+      "eos_token_id": 2,
+      "exponential_decay_length_penalty": null,
+      "finetuning_task": null,
+      "forced_bos_token_id": null,
+      "forced_eos_token_id": null,
+      "hidden_act": "silu",
+      "hidden_size": 4096,
+      "id2label": {
+        "0": "LABEL_0",
+        "1": "LABEL_1"
+      },
+      "initializer_range": 0.02,
+      "intermediate_size": 11008,
+      "is_decoder": false,
+      "is_encoder_decoder": false,
+      "label2id": {
+        "LABEL_0": 0,
+        "LABEL_1": 1
+      },
+      "length_penalty": 1.0,
+      "max_length": 20,
+      "max_position_embeddings": 2048,
+      "min_length": 0,
+      "mlp_bias": false,
+      "model_type": "llama",
+      "no_repeat_ngram_size": 0,
+      "num_attention_heads": 32,
+      "num_beam_groups": 1,
+      "num_beams": 1,
+      "num_hidden_layers": 32,
+      "num_key_value_heads": 32,
+      "num_return_sequences": 1,
+      "output_attentions": false,
+      "output_hidden_states": false,
+      "output_scores": false,
+      "pad_token_id": null,
+      "prefix": null,
+      "pretraining_tp": 1,
+      "problem_type": null,
+      "pruned_heads": {},
+      "remove_invalid_values": false,
+      "repetition_penalty": 1.0,
+      "return_dict": true,
+      "return_dict_in_generate": false,
+      "rms_norm_eps": 1e-06,
+      "rope_scaling": null,
+      "rope_theta": 10000.0,
+      "sep_token_id": null,
+      "suppress_tokens": null,
+      "task_specific_params": null,
+      "temperature": 1.0,
+      "tf_legacy_loss": false,
+      "tie_encoder_decoder": false,
+      "tie_word_embeddings": false,
+      "tokenizer_class": null,
+      "top_k": 50,
+      "top_p": 1.0,
+      "torch_dtype": null,
+      "torchscript": false,
+      "typical_p": 1.0,
+      "use_bfloat16": false,
+      "use_cache": true,
+      "vocab_size": 32000
+    },
+    "tf_legacy_loss": false,
+    "tie_encoder_decoder": false,
+    "tie_word_embeddings": true,
+    "timm_model_ids": [
+      "vit_large_patch14_reg4_dinov2.lvd142m",
+      "vit_so400m_patch14_siglip_224"
+    ],
+    "timm_override_act_layers": [
+      null,
+      null
+    ],
+    "tokenizer_class": null,
+    "top_k": 50,
+    "top_p": 1.0,
+    "torch_dtype": "bfloat16",
+    "torchscript": false,
+    "typical_p": 1.0,
+    "use_bfloat16": false,
+    "use_fused_vision_backbone": true,
+    "vision_backbone_id": "dinosiglip-vit-so-224px"
+  },
+  "rotation_components": 9,
+  "seperate_control_proj": true,
+  "timestep_proj_config": {
+    "num_tokens": 3,
+    "pos_embed_scale": 8,
+    "proj_layers": [
+      128,
+      512,
+      1024
+    ],
+    "time_delta_sec": 0.1
+  },
+  "token_proj_config": {
+    "control_tokens_layers": [
+      4096,
+      2048,
+      1024
+    ],
+    "image_tokens_mode": "vit",
+    "llm_image_tokens_layers": [],
+    "vit_tokens_layers": [
+      2176,
+      1024
+    ]
+  },
+  "token_size": 1024,
+  "transformer_config": {
+    "decoder_block_config": {
+      "dropout": 0.0,
+      "feature_size": 1024,
+      "head_dim": 64,
+      "num_heads": 16
+    },
+    "encoder_block_config": {
+      "feature_size": 1024,
+      "head_dim": 64,
+      "num_heads": 16
+    },
+    "num_blocks": 2,
+    "pos_embed_config": {
+      "embedding_dim": 1024,
+      "num_embeddings": 300
+    }
+  },
+  "transformers_version": "4.44.2"
+}

prismatic_config.py ADDED Viewed

	@@ -0,0 +1,309 @@

+"""
+configuration_prismatic.py
+HuggingFace-style configuration definition for Prismatic VLMs, inheriting from `transformers.PretrainedConfig`.
+Default configuration specifies `siglip-224px+7b`.
+"""
+from typing import Any, Dict, List, Optional
+import transformers
+from transformers import PretrainedConfig
+from transformers.models.auto import CONFIG_MAPPING
+import numpy as np
+# === Utilities for Mapping Prismatic names to HF names ===
+# fmt: off
+VISION_BACKBONE_TO_RESOLUTION: Dict[str, List[int]] = {
+    "clip-vit-l": [224], "siglip-vit-so400m": [224], "dinov2-vit-l": [224], "in1k-vit-l": [224],
+    "clip-vit-l-336px": [336],
+    "siglip-vit-so400m-384px": [384],
+    "dinoclip-vit-l-336px": [336, 336],
+    "dinosiglip-vit-so-224px": [224, 224],
+    "dinosiglip-vit-so-384px": [384, 384],
+}
+VISION_BACKBONE_TO_TIMM_ID: Dict[str, List[str]] = {
+    "clip-vit-l": ["vit_large_patch14_clip_224.openai"],
+    "clip-vit-l-336px": ["vit_large_patch14_clip_336.openai"],
+    "dinov2-vit-l": ["vit_large_patch14_reg4_dinov2.lvd142m"],
+    "in1k-vit-l": ["vit_large_patch16_224.augreg_in21k_ft_in1k"],
+    "siglip-vit-so400m": ["vit_so400m_patch14_siglip_224"],
+    "siglip-vit-so400m-384px": ["vit_so400m_patch14_siglip_384"],
+    "dinoclip-vit-l-336px": ["vit_large_patch14_reg4_dinov2.lvd142m", "vit_large_patch14_clip_336.openai"],
+    "dinosiglip-vit-so-224px": ["vit_large_patch14_reg4_dinov2.lvd142m", "vit_so400m_patch14_siglip_224"],
+    "dinosiglip-vit-so-384px": ["vit_large_patch14_reg4_dinov2.lvd142m", "vit_so400m_patch14_siglip_384"],
+}
+TIMM_OVERRIDE_ACT_LAYER: Dict[str, List[Optional[str]]] = {
+    "clip-vit-l": ["quick_gelu"], "clip-vit-l-336px": ["quick_gelu"],
+    "dinov2-vit-l": [None], "in1k-vit-l": [None],
+    "siglip-vit-so400m": [None], "siglip-vit-so400m-384px": [None],
+    "dinoclip-vit-l-336px": [None, "quick_gelu"],
+    "dinosiglip-vit-so-224px": [None, None], "dinosiglip-vit-so-384px": [None, None]
+}
+LLM_BACKBONE_TO_HF_PATH = {
+    "llama2-7b-pure": "meta-llama/Llama-2-7b-hf", "llama2-13b-pure": "meta-llama/Llama-2-13b-hf",
+    "llama2-7b-chat": "meta-llama/Llama-2-7b-chat-hf", "llama2-13b-chat": "meta-llama/Llama-2-13b-chat-hf",
+    "vicuna-v15-7b": "lmsys/vicuna-7b-v1.5", "vicuna-v15-13b": "lmsys/vicuna-13b-v1.5",
+    "mistral-v0.1-7b-pure": "mistralai/Mistral-7B-v0.1",
+    "mistral-v0.1-7b-instruct": "mistralai/Mistral-7B-Instruct-v0.1",
+    "phi-2-3b": "microsoft/phi-2",
+}
+LLM_BACKBONE_TO_HF_METACLASS = {
+    "llama2-7b-pure": "llama", "llama2-13b-pure": "llama", "llama2-7b-chat": "llama", "llama2-13b-chat": "llama",
+    "vicuna-v15-7b": "llama", "vicuna-v15-13b": "llama",
+    "mistral-v0.1-7b-pure": "mistral", "mistral-v0.1-7b-instruct": "mistral",
+    "phi-2-3b": "phi",
+}
+VALID_VISION_BACKBONES = set(VISION_BACKBONE_TO_RESOLUTION.keys())
+VALID_LLM_BACKBONES = set(LLM_BACKBONE_TO_HF_PATH)
+# fmt: on
+class WaypointTokenizer:
+    """
+    Wraps base LLM/VLM tokenizer and overloads least used token as a control token
+    NOTE: By default, assumes a BPE-style tokenizer akin to the LlamaTokenizer,
+        where *the least used tokens* appear at the end of the vocabulary!
+    TODO: Adding new token vs overloading? When I call `tokenizer.add_token()` vocab stays the same
+    """
+    model_type = "waypointer"
+    is_composition: bool = True
+    def __init__(self, tokenizer: transformers.PreTrainedTokenizerBase, num_tokens: int = 10) -> None:
+        self.tokenizer = tokenizer
+        self.num_tokens = num_tokens
+    def __call__(self, *_) -> str:
+        """Get the text token for control"""
+        return self.tokenizer.decode(self.control_token_ids)
+    @property
+    def control_token_ids(self) -> np.ndarray:
+        # Assumes we're overwriting the final tokens of the vocabulary (least used tokens)
+        return np.arange(self.num_tokens) + int(self.tokenizer.vocab_size - self.num_tokens)
+    @property
+    def num_control_tokens(self) -> int:
+        return self.num_tokens
+class PrismaticConfig(PretrainedConfig):
+    model_type: str = "prismatic"
+    is_composition: bool = False
+    def __init__(
+        self,
+        vision_backbone_id: str = "dinosiglip-vit-so-224px",
+        llm_backbone_id: str = "llama2-7b-pure",
+        arch_specifier: str = "no-align+gelu-mlp", ## TODO: check
+        use_fused_vision_backbone: Optional[bool] = None, ## TODO: check
+        image_resize_strategy: str = "letterbox",
+        text_config: Optional[Dict[str, Any]] = None,
+        llm_max_length: int = 2048,
+        pad_token_id: int = 32000,
+        pad_to_multiple_of: int = 64,
+        output_projector_states: bool = False,
+        **kwargs: str,
+    ) -> None:
+        if vision_backbone_id not in VALID_VISION_BACKBONES:
+            raise ValueError(f"Vision backbone `{vision_backbone_id}` not in {VALID_VISION_BACKBONES = }")
+        if llm_backbone_id not in VALID_LLM_BACKBONES:
+            raise ValueError(f"LLM backbone `{llm_backbone_id}` not in {VALID_LLM_BACKBONES = }")
+        # Set Prismatic Configuration Fields
+        self.vision_backbone_id = vision_backbone_id
+        self.llm_backbone_id = llm_backbone_id
+        self.arch_specifier = arch_specifier
+        self.output_projector_states = output_projector_states
+        # [Contract] All vision backbone parameters are lists =>> supports fused backbones with different preprocessing
+        self.use_fused_vision_backbone = (
+            use_fused_vision_backbone
+            if use_fused_vision_backbone is not None
+            else any(self.vision_backbone_id.startswith(v) for v in ["dinoclip", "dinosiglip"])
+        )
+        self.timm_model_ids = VISION_BACKBONE_TO_TIMM_ID[self.vision_backbone_id]
+        self.timm_override_act_layers = TIMM_OVERRIDE_ACT_LAYER[self.vision_backbone_id]
+        self.image_sizes = VISION_BACKBONE_TO_RESOLUTION[self.vision_backbone_id]
+        self.image_resize_strategy = image_resize_strategy
+        self.hf_llm_id = LLM_BACKBONE_TO_HF_PATH[self.llm_backbone_id]
+        self.llm_max_length = llm_max_length
+        self.pad_token_id, self.pad_to_multiple_of = pad_token_id, pad_to_multiple_of
+        # [IMPORTANT] HF Utilities actually look for a `text_config` field... we need to use that specific naming!
+        self.text_config = (
+            CONFIG_MAPPING[LLM_BACKBONE_TO_HF_METACLASS[self.llm_backbone_id]](**text_config)
+            if text_config is not None
+            else CONFIG_MAPPING[LLM_BACKBONE_TO_HF_METACLASS[self.llm_backbone_id]]()
+        )
+        # Dispatch **kwargs to super() =>> note that `pad_token_id` collides, so we pass it in here as well...
+        super().__init__(pad_token_id=pad_token_id, **kwargs)
+# Here  we need trajectory_vla config, with
+# prismatic_config fields and then the waypointer fields
+class TrajectoryVLAConfig(PretrainedConfig):
+    model_type: str = "trajectoryvla"
+    is_composition: bool = True
+    def __init__(
+        self,
+        prismatic_config = {},
+        token_size: int = 1024,  # Timestep token size
+        cheat: bool = False,  # If True, cheat and use action tokens; Works only with OpenVLA checkpoint
+        num_timesteps: int = 20,  # Number of prediction time steps
+        rotation_components: int = 9,  # Number of rotation componens: euler -> 3, quaternion -> 4, rotmat -> 9
+        num_timestep_tokens : int = 3,
+        seperate_control_proj: bool = True,  # If True, project control components separately
+        timestep_proj_config: Dict[str, Any] = {},
+        token_proj_config: Dict[str, Any] = {},
+        transformer_config: Dict[str, Any] = {},
+        # prismatic_config: PrismaticConfig,
+        # waypointer_config: Dict[str, Any],
+        # **kwargs: str,
+    ):
+        # super().__init__(**prismatic_config)
+        self.prismatic_config = PrismaticConfig(**prismatic_config)
+        self.token_size = token_size
+        self.cheat = cheat
+        self.num_timesteps = num_timesteps
+        self.rotation_components = rotation_components
+        self.seperate_control_proj = seperate_control_proj
+        self.timestep_proj_config = timestep_proj_config
+        self.token_proj_config = token_proj_config
+        self.transformer_config = transformer_config
+        # self.num_timestep_tokens = num_timestep_tokens
+    @property
+    def control_components(self) -> int:
+        # Number of control dimensions: 3 translation, N rotation, 1 gripper
+        return 3 + self.rotation_components + 1
+    @property
+    def num_timestep_tokens(self) -> int:
+        return self.timestep_proj_config['num_tokens']
+# class WaypointerConfig(ConfigurableModuleConfig):
+#     token_size: int = 1024  # Timestep token size
+#     cheat: bool  # If True, cheat and use action tokens; Works only with OpenVLA checkpoint
+#     timestep_proj_config: AutoConfig  # Timestep tokens
+#     token_proj_config: TokenProjectorConfig  # LLM output tokens projection and packing
+#     transformer_config: AutoConfig  # Transformer config
+#     # Output configurations
+#     num_timesteps: int = 20  # Number of prediction time steps
+#     rotation_components: int = 3  # Number of rotation componens: euler -> 3, quaternion -> 4, rotmat -> 9
+#     separate_control_proj: bool = True  # If True, project control components separately
+#     @property
+#     def control_components(self) -> int:
+#         # Number of control dimensions: 3 translation, N rotation, 1 gripper
+#         return 3 + self.rotation_components + 1
+#     @property
+#     def num_timestep_tokens(self) -> int:
+#         return self.timestep_proj_config.num_tokens
+class OpenVLAConfig(PrismaticConfig):
+    model_type: str = "openvla"
+    def __init__(
+        self,
+        norm_stats: Optional[Dict[str, Dict[str, Dict[str, Dict[str, List[float]]]]]] = None,
+        n_action_bins: int = 256,
+        **kwargs: str,
+    ) -> None:
+        self.norm_stats, self.n_action_bins = norm_stats, n_action_bins
+        super().__init__(**kwargs)
+if  __name__ == "__main__" :
+    # yaml_file = 'barrel/pipes/vlams/configs/waypoints/waypointer_multistep_fractal.yaml'
+    prismatic_config = PrismaticConfig()
+    print(prismatic_config)
+    prismatic_config_dict = {
+        "vision_backbone_id":"dinosiglip-vit-so-224px",
+        # "llm_backbone_id":"llama2-7b-pure",meta-llama/Llama-2-7b-hf
+        "llm_backbone_id": "meta-llama/Llama-2-7b-hf",
+        "arch_specifier": "no-align+gelu-mlp", ## TODO: check
+        "use_fused_vision_backbone" :None, ## TODO: check
+        "image_resize_strategy" : "letterbox",
+        "text_config" : None,
+        "llm_max_length"  : 2048,
+        "pad_token_id" :32000,
+        "pad_to_multiple_of" : 64,
+        "output_projector_states" : False,
+    }
+    token_proj_config = {
+        "vit_tokens_layers": [2176, 1024],
+        "control_tokens_layers": [4096, 2048, 1024],
+        "image_tokens_mode": 'vit',
+    }
+    timestep_proj_config = {
+        "pos_embed_scale": 1.0,
+        "proj_layers": [1024],
+        "time_delta_sec": 0.1,
+        "num_tokens":3
+    }
+    TrajectoryVlaConfig = {
+        "prismatic_config":prismatic_config_dict,
+        "token_size": 1024,
+        "cheat": False,
+        "num_timesteps": 20,
+        "rotation_components": 3,
+        "seperate_control_proj": True,
+        "timestep_proj_config": {},
+        "token_proj_config": {},
+        "transformer_config": {},
+    }
+    TrajectoryVLAConfig = TrajectoryVLAConfig( **TrajectoryVlaConfig)
+    print(TrajectoryVLAConfig)
+class WaypointTokenizer:
+    """
+    Wraps base LLM/VLM tokenizer and overloads least used token as a control token
+    NOTE: By default, assumes a BPE-style tokenizer akin to the LlamaTokenizer,
+        where *the least used tokens* appear at the end of the vocabulary!
+    TODO: Adding new token vs overloading? When I call `tokenizer.add_token()` vocab stays the same
+    """
+    def __init__(self, tokenizer: transformers.PreTrainedTokenizerBase, num_tokens: int = 10) -> None:
+        self.tokenizer = tokenizer
+        self.num_tokens = num_tokens
+    def __call__(self, *_) -> str:
+        """Get the text token for control"""
+        return self.tokenizer.decode(self.control_token_ids)
+    @property
+    def control_token_ids(self) -> np.ndarray:
+        # Assumes we're overwriting the final tokens of the vocabulary (least used tokens)
+        return np.arange(self.num_tokens) + int(self.tokenizer.vocab_size - self.num_tokens)
+    @property
+    def num_control_tokens(self) -> int:
+        return self.num_tokens