Research and build efficient ML systems for large-scale LLMs and agentic RL: design algorithms and system techniques, prototype in training/inference stacks, run large-scale experiments, and translate findings into production or publications.
About Goaly
About the Role
Core Responsibilities
Requirements
Bonus Points
At Goaly, our mission is to make custom AI affordable for every business. Our founding team comes from the front lines of top AI labs and tech giants (Meta MSL, TikTok AI, Google DeepMind, xAI, Microsoft Research, etc.), where we built large-scale training infrastructure powering trillion-parameter models and scaled GenAI models to a global user base. Now, we are building something we wish we had before: a platform that makes training and adapting custom AI affordable for all modern companies, not just Big Tech. Our north star is ambitious: for a domain-specific task, reach 90% of SOTA performance at less than 10% of the cost. To get a taste of what we are doing, see our first tech blog.
As an AI Research Scientist (Efficient ML Systems) at Goaly, you will research and build the systems that make frontier-scale models practical. This role sits at the intersection of algorithms, systems, and hardware efficiency.You will design and evaluate new training and inference techniques, prototype them in real systems, and push them to production-scale workloads.
Your work will either ship directly into our core platform or lead to publications at top venues such as NeurIPS, ICML, ICLR, or CVPR.This is not a paper-only role. You will write real systems code, run large-scale experiments, and directly shape how modern LLMs and RL systems are trained and deployed.
- Research efficient ML systems: Invent and evaluate algorithms and system techniques that improve LLM and agentic RL training and inference efficiency (memory, compute, communication, and stability).
- Scale agentic RL: Design and optimize large-scale agentic RL pipelines, including asynchronous training, experience management, reward modeling, and long-horizon stability.
- End-to-end experimentation: Design large-scale experiments spanning model architecture, training algorithms, distributed systems, and hardware-aware optimization.
- System-aware research: Prototype research ideas directly in training and inference stacks (e.g., parallelism strategies, attention kernels, RL training pipelines) and validate them at scale.
- Production & publication: Translate successful ideas into production-ready systems and/or publish them at top-tier conferences with full internal support.
- Ph.D. or Master's degree in CS, AI, Systems, or related fields (Exceptional undergraduates with strong research capabilities may be considered).
- Strong foundation in LLM or large-scale ML training, including Transformers, attention mechanisms, distributed training, and optimization methods.
- Experience or strong interest in agentic RL or large-scale reinforcement learning systems, including stability, scalability, or long-horizon training challenges.
- Demonstrated interest in efficiency-focused research, such as training acceleration, memory optimization, parallelism, kernels, or RL system robustness.
- Proficient in PyTorch or JAX. Clean coding style and strong command of Python.
- Adaptability: A fast learner with a strong sense of responsibility, capable of wearing multiple hats and handling cross-stack challenges.
- First-author publications at top conferences (NeurIPS, ICML, ICLR, CVPR, ACL).
- High-star open-source projects on Hugging Face or Gold/Silver medals in Kaggle competitions.
Similar Jobs
Cloud • Insurance • Payments • Software • Business Intelligence • App development • Big Data Analytics
Lead UX research and design for B2B/SaaS insurance products: create wireframes, mockups, and prototypes; run user research and usability tests; use analytics to measure outcomes; collaborate with product and engineering to implement consistent, validated UX solutions.
Top Skills:
BalsamiqFigmaMiroWhiteboards
Cloud • Insurance • Payments • Software • Business Intelligence • App development • Big Data Analytics
Design, build, and operate enterprise-scale multi-cloud infrastructure (Azure primary, GCP, AWS exposure). Own landing zones, Terraform modules, production AKS/GKE Kubernetes, Vault secrets, hybrid networking, CI/CD pipelines, monitoring, DR, and automation (Ansible, Python/Bash). Mentor engineers, document runbooks, and collaborate with security, application teams, and leadership to ensure secure, reliable, cost-optimized cloud platforms.
Top Skills:
AksAnsibleApp GatewayArtifact RegistryAWSAwxAzureAzure DevopsAzure MonitorAzure StorageBashBgpBigQueryCloud BuildCloud LoggingCloud RunCloud SqlCloudboltDatadogDnsEc2EksGitlab CiGkeGoogle Cloud MonitoringGoogle Cloud Platform (Gcp)Hashicorp VaultHelmIamJenkinsKubernetesLoad BalancingManaged IdentityNsgPowershellPrivate EndpointsPythonS3SignozTerraformVertex AiVpcVpc Service ControlsVpnWorkload Identity
Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Lead finance transformation engagements using Oracle Cloud ERP and EPM. Design and implement Oracle Financials and Hyperion solutions, integrate RPA/ML/analytics, ensure compliance, manage stakeholder relationships, coach teams, and drive strategic outcomes on large, cross-border projects.
Top Skills:
Ahcs/FahAnalyticsFixed Assets (Fa)Hyperion Financial ManagementMachine LearningOracle ApOracle ArOracle Cloud ErpOracle CmOracle EpmOracle ExpensesOracle FinancialsOracle GlOracle Ppm (Grants)Project BillingProject CostingRpa
What you need to know about the Boston Tech Scene
Boston is a powerhouse for technology innovation thanks to world-class research universities like MIT and Harvard and a robust pipeline of venture capital investment. Host to the first telephone call and one of the first general-purpose computers ever put into use, Boston is now a hub for biotechnology, robotics and artificial intelligence — though it’s also home to several B2B software giants. So it’s no surprise that the city consistently ranks among the greatest startup ecosystems in the world.
Key Facts About Boston Tech
- Number of Tech Workers: 269,000; 9.4% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Thermo Fisher Scientific, Toast, Klaviyo, HubSpot, DraftKings
- Key Industries: Artificial intelligence, biotechnology, robotics, software, aerospace
- Funding Landscape: $15.7 billion in venture capital funding in 2024 (Pitchbook)
- Notable Investors: Summit Partners, Volition Capital, Bain Capital Ventures, MassVentures, Highland Capital Partners
- Research Centers and Universities: MIT, Harvard University, Boston College, Tufts University, Boston University, Northeastern University, Smithsonian Astrophysical Observatory, National Bureau of Economic Research, Broad Institute, Lowell Center for Space Science & Technology, National Emerging Infectious Diseases Laboratories

.png)
