Applied AI Research Engineer

Appen 2

Sign In to Apply

Pay not listed

  • Remote
  • Full-time
  • dia
  • AI Data Opportunities
  • 5d ago

Job Description

About the Role 

As an Applied Research Engineer, you’ll build practical AI research assets that support Frontier lab initiatives and customer engagements. This is an implementation-focused role for someone who enjoys turning research concepts into working systems. 

You’ll work with a high degree of autonomy, experimenting with new approaches and developing solutions that can be reused across customer opportunities. You’ll partner closely with the GenAI Research team and cross-functional stakeholders to bring technical ideas into practical applications. 

Your Impact 

  • Build reinforcement learning and agent environments for real customer and Frontier lab use cases, including task specifications, scoring, and evaluation.
  • Develop benchmarks and evaluation harnesses to measure model and data quality across areas such as accuracy, robustness, safety, latency, and cost.
  • Build LLM pipelines and agentic systems that support research, evaluation, and customer trials.
  • Run fine-tuning, adapter, and other model experiments to evaluate how data and methods influence model behavior.
  • Deploy local or self-hosted models for evaluation, inference, and automation workflows.
  • Document experiments, configurations, data, results, and known limitations so other engineers can reproduce and build on your work.
  • Partner with the GenAI Research team and cross-functional stakeholders to turn technical work into reusable assets for customer engagements. 

What You Bring 

  • Bachelor’s, Master’s, or PhD in Computer Science, Engineering, Machine Learning, or a related technical field.
  • 3+ years of professional engineering or relevant industry experience in AI/ML or software engineering.
  • Strong software engineering skills and experience building reliable, maintainable AI systems.
  • Hands-on experience building agentic systems, reinforcement learning environments, LLM pipelines, or similar AI systems.
  • Experience building evaluation harnesses, benchmarks, or model testing pipelines.
  • Ability to work independently on technical problems and move quickly from an idea or research question to a working solution.
  • Strong understanding of experimentation, reproducibility, and technical documentation. 

Nice to Haves 

  • Developed synthetic data generation systems or datasets.
  • Published research papers, benchmarks, or other technical research.
  • Worked with SWE-bench or similar software engineering evaluation environments.
  • Built or deployed local inference, open-weight models, or self-hosted model environments.