About Nscale
Nscale is taking on the hyperscalers by building a vertically integrated GenAI cloud platform, owning the data centres, software, and applications that power today's AI stack.
About the Role
Nscale is looking for Senior / Staff AI Engineers to join the core AI team and build the systems that power the GenAI cloud platform, designing and optimising distributed systems for large-scale training, post-training, evaluation, and low-latency, high-throughput inference.
Responsibilities
- Drive inference performance and efficiency, including KV cache management, continuous batching, speculative decoding, and quantisation.
- Build and improve post-training services, including fine-tuning (LoRA, QLoRA), alignment (RLHF, DPO), and agentic RL.
- Develop evaluation and benchmarking systems to measure model quality, safety, and system performance.
- Build developer-facing APIs, SDKs, and tooling that enable other engineers to use Nscale's AI services.
Requirements
- 5+ years of experience building production systems in machine learning, distributed systems, or high-performance infrastructure.
- Deep understanding of transformer architectures, LLMs, and/or multimodal models.
- Strong proficiency in Python and PyTorch, with production-grade ML systems experience.
- Experience with GPU/accelerator optimisation (CUDA, ROCm) and distributed compute paradigms.