About Nscale
Nscale is taking on the hyperscalers by building a vertically integrated GenAI cloud platform, owning the data centres, software, and applications that power today's AI stack.
About the Role
As a Senior Technical Product Manager for Observability, you own the platform that gives customers and internal operators real-time visibility into their GPU fleet: the telemetry pipeline, the aggregation and storage layer, and the observability surfaces that enable fleet management, incident response, and alerting at scale.
What You'll Be Doing
- Own the roadmap for Nscale's observability platform: telemetry pipeline, log and metrics aggregation, trace collection, and customer facing dashboards.
- Define how logs, metrics, and traces are captured from physical infrastructure and surfaced to enable customers to manage their fleet.
- Own alerting strategy and optimisation, reducing noise and ensuring the right signal reaches the right person.
- Define and drive metrics that matter: alert signal-to-noise ratio, time-to-detect, time-to-resolve, telemetry coverage.
What You Need
- 5–8 years in product management, with a track record owning significant areas in observability, infrastructure, or operations-facing products.
- Hands-on experience with Prometheus, Loki, Mimir, Datadog, Grafana, or OpenTelemetry.
- Experience with deployment tooling in a data centre or infrastructure context.