Embodied AI / VLA Research Engineer

Maxinsights

Sign In to Apply

Pay not listed

  • On-site
  • Full-time
  • Santa Clara
  • ML & Data
  • 1d ago

Job Description

Job Description:

Position Overview

In this role, you will work on the research, training, optimization, and real-world deployment of embodied AI models, including Vision-Language-Action (VLA) models, World Models, and related robotics foundation models.

You will work across multimodal perception, robot control, long-horizon task planning, large-scale robot data, and model deployment. This is a highly hands-on role that involves both model development and real-world robotic system integration.

Key Responsibilities

Embodied AI Model Development

  • Research and develop Vision-Language-Action (VLA), World Model, and other embodied AI foundation models.

  • Work on areas including:

    • Robotic manipulation

    • Multimodal perception and control

    • Long-horizon task planning

    • Vision-language-action reasoning

    • Robot-environment interaction

  • Develop and optimize models for real-world robotic applications.

  • Translate research ideas into practical models and systems that can operate reliably on physical robots.

Foundation Model Training & Optimization

  • Train and optimize embodied AI foundation models using large-scale real-robot datasets and egocentric human demonstration data.

  • Design approaches for incorporating multimodal inputs such as:

    • Vision

    • Force

    • Tactile sensing

    • Proprioception

    • Other robot and environmental signals

  • Develop and evaluate multimodal fusion architectures.

  • Conduct model training, evaluation, benchmarking, and performance optimization.

  • Analyze model performance and identify opportunities to improve training efficiency, generalization, and real-world performance.

Robotics Data Pipeline

  • Build and improve large-scale robotics data pipelines for model training.

  • Develop processes for:

    • Data cleaning

    • Data filtering

    • Resampling

    • Data augmentation

    • Data quality evaluation

  • Design scalable data pipelines capable of supporting large volumes of robot and human demonstration data.

  • Work closely with data and robotics teams to improve dataset quality and training efficiency.

Model Deployment & Real-Robot Testing

  • Deploy trained models to physical robot systems.

  • Perform real-world robot debugging, testing, and performance optimization.

  • Diagnose issues across models, sensors, software, and robotic hardware.

  • Iterate between model training and real-world testing to improve system performance.

  • Help ensure models operate reliably and consistently in real-world environments.

Qualifications

Required

  • 1+ years of relevant industry or research experience in machine learning, robotics, computer vision, embodied AI, or a related field.

  • Strong understanding of deep learning and modern machine learning methods.

  • Experience with PyTorch or similar deep learning frameworks.

  • Experience training and evaluating machine learning models.

  • Strong programming skills in Python and familiarity with relevant ML/robotics tooling.

  • Understanding of multimodal learning, computer vision, robotics, or related areas.

  • Ability to work in a fast-paced startup environment and take ownership of technical problems from research through implementation.

  • Strong problem-solving and debugging skills.

Strong Plus

Real-World Robotics Deployment

  • Experience deploying and debugging machine learning models on physical robot systems.

  • Ability to bring models from development into real-world robotic environments.

  • Experience troubleshooting and stabilizing robotic systems in production or experimental environments.

Multimodal / VLA Models

  • Experience working with force, tactile, or other multimodal sensing.

  • Experience designing or training multimodal fusion models.

  • Hands-on experience with VLA models, including model design, training, evaluation, or real-world applications.

  • Experience with robotic manipulation or embodied AI systems.

Large-Scale Distributed Training

  • Experience with large-scale distributed model training.

  • Familiarity with DDP, DeepSpeed, FSDP, or similar distributed training frameworks.

  • Experience optimizing training performance, GPU utilization, memory usage, or training throughput.

  • Experience working with large-scale datasets and distributed data pipelines.

Ideal Candidate

We are looking for an engineer who is excited about the intersection of foundation models and physical robotics.

The ideal candidate is:

  • Hands-on and comfortable moving between research, coding, experimentation, and real-world robot testing.

  • Interested in solving problems that cannot be addressed through simulation or software alone.

  • Comfortable working with large-scale datasets and modern foundation-model architectures.

  • Able to take ownership of a problem from data → training → evaluation → deployment → real-world iteration.

  • Comfortable working in an early-stage environment where priorities can move quickly.

  • Curious about emerging VLA, World Model, and embodied AI research and able to translate new ideas into working systems.

Why Join MaxInsights?

  • Work directly on embodied AI and robotics foundation models.

  • Work with large-scale real-world robotics and human demonstration data.

  • Gain hands-on experience across the full AI development lifecycle, from data pipelines to real-robot deployment.

  • Work in a fast-moving startup environment with significant ownership and technical autonomy.

  • Collaborate with teams working at the forefront of robotics and foundation-model development.

Default Benefits:

  • Health insurance

  • Vision care

  • Dental coverage

  • 401(k)

  • Paid holidays

  • PTO (Paid Time Off)

  • Sick leave