Phaidra Senior AI Research Engineer at Phaidra owning research infrastructure and performance engineering for industrial AI control systems. Responsible for end-to-end research-to-production lifecycle.
Responsibilities
Wear different hats across the research-to-production lifecycle: ML Engineer, ML-Ops Engineer, Software Engineer, and Performance Engineer.
Own and evolve our research infrastructure end-to-end, from experiment orchestration and distributed training to model tracking, evaluation, and automated deployment, so researchers can move from idea to validated result quickly.
Build and scale distributed compute for research workloads (e.g. Ray-based training and data pipelines on Kubernetes/GCP), including managing GPU capacity across zones/regions and keeping experiment infrastructure reliable and cost-efficient.
Improve the speed and quality of our R&D through performance engineering: vectorizing and parallelizing simulators and training code, profiling bottlenecks, and driving large speedups.
Deeply understand the capabilities and tools offered by Phaidra’s internal platform and how to utilize them to best serve our customers.
Maintain clear and concise documentation of your research, products and actions.
Participate in making decisions for the medium-to-long-term vision impacting Research and Phaidra.
Mentor peers and delegate tasks within the team, owning the project delivery.
Act as a point of contact between Research and Production engineering teams to productionize new breakthroughs rapidly.
Key Qualifications
4+ years of progressive relevant work experience after Master’s graduation.
6+ years of progressive relevant work experience after Bachelor’s graduation.
Qualification
Exposure to reinforcement learningPyTorchDockerIn your first 30 days
Preferred
Understanding of industrial heating and cooling processes and their applications within manufacturing or data center environments.
Previous research experience in the field of ML or AI or MLOps experience.
Experience with Ray for distributed computing and orchestrating multi-node GPU workloads.
Exposure to reinforcement learning, simulation, and control systems.
A general scientific or physical-sciences foundation that helps you collaborate closely with researchers.