hirly

Apply with hirly

Senior AI Engineer (Kubernetes & Customised Scheduler)

Firmus Technologies · Singapore

Upload your resume to see how well you match this job — free, in seconds, no account needed.

Your resume is used only to score it against this job. If you don't create an account, it is deleted within 24 hours.

Firmus Technologies Firmus Technologies is a global leader pioneering the development and operation of efficient AI infrastructure across Asia Pacific. Founded in Australia in 2019, our mission is to create the most efficient AI infrastructure by combining cutting-edge technology with a steadfast commitment to sustainability. At Firmus, we are unique in our approach. We design, build, and operate a new class of digital infrastructure – the AI Factory. Through our model-to-grid technology approach, we have pushed the boundaries of multi-generational liquid cooling systems, energy management, AI software orchestration, and construction. For our customers, this approach allows us to make every watt count and deliver low-cost AI tokens globally. Firmus AI Cloud Our large-scale GPU cloud platform, Firmus AI Cloud, is purpose-built to deliver energy-efficient AI compute at scale to customers. It empowers developers, enterprises, educational institutions, and government users to train and deploy AI models with unmatched efficiency and cost savings. With an ever-growing suite of services and applications, we are committed to delivering a cloud experience that is market-leading, proprietary, and built to scale. Why Firmus? As an NVIDIA Cloud and Engineering partner in Asia Pacific, you will gain skills, experience, and exposure across the AI industry and be part of shaping what this industry looks like for decades to come. We are founder-led, not a big corporate. Decisions happen fast, our leaders are accessible, and there's minimum bureaucracy between you and the work. Ownership comes early. Whatever your role, you will have a direct line to outcomes, helping shape how the business grows as we scale nationally across a long-term, large-scale roadmap. Work alongside founders and experts in AI infrastructure, energy systems and next-generation compute. What we build here has impact beyond the business. Our AI Factories are designed to operate as assets to the energy grid to actively strengthen the communities and regions they operate in rather than drawing from them. Considering applying? You don't need a perfect background to join our team. If you're driven and curious, there's a path for you. We back our people to grow into new domains and take on challenges beyond their previous experience.

Role Summary

The Senior AI Engineer will be a core builder of the AI & Applications team’s Model-to-Grid product, designing, developing, and operating the Kubernetes-based workload orchestration layer for massive-scale AI factories. The role has a strong software and platform engineering focus: building a proprietary job scheduler that enables efficient, reliable, and policy-driven execution of training, fine-tuning, inference, batch, benchmarking, and agentic workloads across cutting-edge GPU systems, including NVIDIA NVL72 GB300-scale deployments and future-generation platforms such as VR200. The proprietary scheduler is a central product differentiator. It must be network-topology aware and AI-factory-resource aware: making placement, queueing, prioritization, admission, and execution decisions based not only on nominal GPU availability, but also on GPU and NVLink/NVSwitch topology, node and fault-domain boundaries, RDMA and fabric health, storage locality and throughput, workload characteristics, capacity, power, thermal state, maintenance activity, and other operational constraints. The resulting capability improves outcomes in two directions. For AI users, it provides better workload placement, lower queue times, higher GPU utilization, stronger job-success rates, improved end-to-end throughput, and faster time-to-results. For AI-factory and grid operators, it serves as a technical shock absorber by making demand more observable, controllable, schedulable, and responsive to infrastructure availability, system health, power, thermal, capacity, and operational conditions. The role will work closely with other AI engineers, inference and optimization engineers, the Model-to-Grid product and program lead, Platform, Infrastructure, Security, and operations teams. It will turn scheduler product requirements into robust production capabilities, integrating Kubernetes, custom controllers, APIs, observability, automation, workload recipes, benchmark signals, and supporting platform services. The work will align with the architectural direction of

Nvidia Dsx Os

and AI Factory Blueprint concepts

  • co-designed, resilient, multi-tenant AI-factory operations, while delivering the organization’s proprietary scheduler intelligence and differentiated Model-to-Grid capabilities.

Key Responsibilities

  • Design, build, operate, and continuously improve the proprietary Kubernetes-native job scheduling platform for AI workloads.
  • Develop scheduler architecture, custom resource definitions, Kubernetes controllers, admission webhooks, scheduling plugins, APIs, CLI tools, and automation required to support workload submission, placement, execution, monitoring, and recovery.
  • Define and implement scheduling policies for training, fine-tuning, inference, benchmarking, batch, data-processing, and agentic workloads.
  • Build topology-aware placement mechanisms that account for GPU locality, NVLink/NVSwitch domains, node topology, NUMA affinity, NIC placement, RDMA paths, network-fabric topology, storage locality, and fault-domain boundaries.
  • Build AI-factory resource-aware scheduling mechanisms that account for cluster capacity, node health, fabric condition, storage performance, GPU availability, maintenance windows, software compatibility, capacity reservations, power limits, thermal conditions, and operational constraints.
  • Implement workload-control capabilities including admission control, queueing, priorities, quotas, fair sharing, reservations, preemption, gang scheduling, co-scheduling, backfilling, workload aging, retry policies, checkpoint-aware scheduling, deferred execution, and failure recovery.
  • Develop workload demand-shaping and grid-aware controls that allow eligible workloads to be delayed, paced, prioritized, rescheduled, right-sized, or placed differently in response to capacity, health, power, thermal, maintenance, or other AI-factory operating signals.
  • Integrate the scheduler with Kubernetes control-plane capabilities, GPU device plugins, node-feature discovery, network and storage services, observability systems, and policy engines.
  • Integrate, where appropriate, with Slurm, Slinky, vCluster, Kueue, KAI, Volcano, YuniKorn, or related systems to deliver a coherent workload-management experience across AI and HPC use cases.
  • Build a stable developer and user experience for workload submission and management, including APIs, SDKs, CLI workflows, templates, workload definitions, job-status visibility, event streams, scheduling explanations, and self-service troubleshooting capabilities.
  • Work with AI and inference engineers to encode validated model and workload recipes into scheduler-aware templates, including requirements for model size, precision, distributed parallelism, GPU count, topology, network, storage, runtime, benchmark target, and expected resource profile.
  • Enable Model-to-Grid benchmarking by exposing scheduler, placement, resource, and workload-lifecycle data for end-to-end analysis across models, runtimes, GPUs, network, storage, power, thermals, and application performance.
  • Build observability into the scheduling platform, including queue depth, scheduling latency, admission outcomes, placement decisions, resource fragmentation, topology quality, GPU utilization, job lifecycle, failure reasons, preemption events, retry behavior, power and thermal signals, and workload performance.
  • Develop actionable scheduler explanations and operator views that show why a workload was queued, admitted, placed, deferred, preempted, or rescheduled and what changes could improve execution outcomes.
  • Establish performance t