Sign up to access all features of our service
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff Slurm Cluster & HPC Engineer

Full-time

Bitdeer Technologies Group

About Bitdeer:

Bitdeer is a world-leading technology company for Bitcoin mining and AI cloud.

Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers. Apart from designing industry-leading ASIC chips and manufacturing mining rigs, the Group handles complex processes involved in computing across the value chain. This includes equipment procurement, transport logistics, datacenter design and construction, equipment management, and network and facility operations. Bitdeer also offers advanced cloud capabilities to customers with a high demand for artificial intelligence.

Headquartered in Singapore, Bitdeer operates globally with a diversified 3 GW energy portfolio, and deploys Bitcoin mining and HPC datacenters in the United States, Bhutan, Norway, Canada, Malaysia, and Ethiopia.

What you will be responsible for:

  • Slurm cluster architecture and lifecycle — Design, deploy, and operate production Slurm clusters on bare metal and VMs: slurmctld/slurmdbd high availability, slurmrestd, configless slurmd, SACK/MUNGE and JWT authentication, and rolling version upgrades on live clusters without losing running jobs.
  • Topology-aware scheduling for GPU fabrics — Model the physical fabric in topology.conf — topology/tree for rail-optimized InfiniBand/RoCE designs and topology/block for NVLink domains such as GB200/GB300 NVL72 — and prove placement quality with NCCL bandwidth and multi-node training validation rather than assumption.
  • Multi-tenant scheduling policy — Own the account/association tree, partitions, QOS, fairshare, preemption, reservations, and per-tenant TRES limits. Enforce fail-closed defaults: an unresolved tenant identity or an empty entitlement set must deny, never degrade into unrestricted access.
  • Slinky on Kubernetes — Lead implementation of the Slinky slurm-operator, including its NodeSet, LoginSet, Accounting, RestAPI, and Token custom resources, cert-manager and Helm-based delivery, shared parallel-storage mounts, and login pods running sackd/sshd. Evaluate and pilot slurm-bridge for co-scheduling Kubernetes Pods, PodGroups, Jobs, JobSets, and LeaderWorkerSets through the Slurm scheduler, and document its constraints — notably exclusive whole-node allocation — before any customer exposure.
  • Elastic capacity between Slurm and Kubernetes — Use Slurm cloud and power-save mechanisms (ResumeProgram/SuspendProgram, SuspendTime, ResumeTimeout) together with fleet automation to shift GPU nodes between batch training queues and Kubernetes inference capacity as demand moves.
  • Container and job runtime — Operate Pyxis/Enroot and OCI/containerd job paths with correct gres.conf, cgroup v2 device constraints, and CUDA_VISIBLE_DEVICES behavior; support MPI/PMIx, module/Spack environments, and customer-supplied images.
  • Cluster health and reliability engineering — Build the passive and active health-check system expected of a top-tier GPU cloud: prolog/epilog checks, LBNL NHC or equivalent, DCGM diagnostics, and detection of XID/SXID errors, ECC faults, PCIe errors, GPUs falling off the bus, IB/RoCE link flaps, and NCCL stalls — with automatic drain and job requeue. Own burn-in and acceptance testing for every new rack before it carries paid work.
  • Automation and infrastructure as code — Deliver clusters through Terraform/Ansible, golden images, and bare-metal provisioning (PXE, Redfish, IPMI) so that a cluster build is reproducible, reviewable, and auditable rather than hand-tuned.
  • Observability, accounting, and billing integration — Instrument queue wait time, allocation efficiency, GPU utilization, and job failure taxonomy through a Slurm exporter into Prometheus/Grafana; configure AccountingStorageTRES and TRESBillingWeights, and reconcile sacct/sreport GPU-hours against the platform's metering and invoicing pipeline.
  • Technical leadership and customer engagement — Write runbooks and tenant-facing documentation, onboard and support enterprise customers, act as escalation point for cluster incidents, and mentor platform engineers on Slurm and HPC scheduling practice.

How you will stand out:

  • 8+ years in HPC, systems, or cloud infrastructure engineering, including 4+ years operating production Slurm clusters at 100+ GPU-node scale with real users and service-level commitments.
  • Deep hands-on Slurm expertise: slurm.conf, gres.conf, topology.conf, cgroup.conf, partitions/QOS/fairshare/preemption/reservations, slurmdbd accounting, slurmrestd, MUNGE/SACK and JWT authentication, and version upgrades performed on live clusters.
  • Strong GPU and fabric fundamentals: NVIDIA drivers and Fabric Manager, DCGM, MIG, InfiniBand/RoCEv2 (subnet manager/UFM, rail-optimized topology), GPUDirect RDMA, and practical NCCL tuning and failure diagnosis.
  • Production Kubernetes experience and working knowledge of the operator/CRD pattern, plus hands-on exposure to at least one Slurm-on-Kubernetes stack — Slinky slurm-operator or slurm-bridge, CoreWeave SUNK, or Nebius Soperator — with an informed view of the tradeoffs between them.
  • Experience delivering both bare-metal and virtualized compute: bare-metal provisioning and firmware/BIOS lifecycle management, hypervisor or VM-based clusters (KVM/QEMU or a public-cloud equivalent), and Terraform/Ansible-driven automation.
  • Working knowledge of parallel and shared storage for AI workloads — Lustre, GPFS/Spectrum Scale, WEKA, VAST, or NFS — and of how storage behavior shapes job performance and failure modes.
  • Proficient in Python and Bash for cluster automation; Go experience is a plus for integrating with Bitdeer AI's platform control plane and with Slurm/Slinky REST client code.
  • Multi-tenant security discipline: derives tenant scope from a verified identity rather than client-supplied fields, designs authorization to fail closed, and treats isolation across accounts, namespaces, storage, and networks as a hard requirement.
  • Clear written and verbal communication in English, with the maturity to work directly with enterprise customers and to translate scheduling and reliability tradeoffs for product, sales, and executive stakeholders.

What you will experience working with us:

  • A culture that values authenticity and diversity of thoughts and backgrounds;
  • An inclusive and respectable environment with open workspaces and exciting start-up spirit;
  • Fast-growing company with the chance to network with industrial pioneers and enthusiasts;
  • Ability to contribute directly and make an impact on the future of the digital asset industry;
  • Involvement in new projects, developing processes/systems;
  • Personal accountability, autonomy, fast growth, and learning opportunities;
  • Attractive welfare benefits and developmental opportunities such as training and mentoring.

--------------------------------------------------------------------

Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, colour, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union. 


#LI-ST1

Vacancy posted 25 days ago
Similar jobs that could be interesting for youBased on the Staff Slurm Cluster & HPC Engineer in Singapore vacancy
  • 5000 - 8000 SGD

     ...Purpose Client is seeking an experienced HPC Systems Engineer (or Senior HPC Systems Engineer,...  ...operate, and maintain Linux-based HPC clusters, including compute, storage, and high-speed...  ...Support: HPC job schedulers (e.g. Slurm, PBS Pro, LSF) Parallel file systems... 

    OPENSOURCE TECHNOLOGIES PTE. LTD.

    Singapore
    8 days ago
  • 6000 - 11500 SGD

     ...leader in Singapore to hire experienced HPC Engineers to support the implementation, configuration...  ...of High-Performance Computing (HPC) clusters, parallel file systems, and associated infrastructure...  ...job schedulers such as PBS, PBS-Pro, and Slurm. Experience configuring, supporting,... 

    XCELLINK PTE. LTD.

    Singapore
    25 days ago
  • 6000 - 8500 SGD

     ...seeking a High-Performance Computational (HPC) Engineer/ Senior HPC Engineer/Scientist (depending...  ...environment used by research scientists, staff and students. Key Responsibilities:...  ...Administration and operation of several HPC Linux clusters, storage, networking and associated... 

    FUJITSU ASIA PTE LTD

    Singapore
    11 days ago
  • 6000 - 9000 SGD

     ...Responsibilities Support day-to-day operations of HPC clusters, including compute nodes, storage systems...  ...management systems and job schedulers (Slurm, PBS, LSF) for batch processing and...  ...modules * Collaborate with senior engineers on cluster optimization, scaling, and performance... 

    D L RESOURCES PTE LTD

    Singapore
    13 hours ago
  •  ...We are looking for HPC Engineers to design, deploy and operate high-performance computing infrastructure...  ...workloads. You will manage compute clusters, parallel storage, high-speed...  ...Administer job scheduling platforms such as Slurm, PBS Pro, or LSF Maintain and optimise... 

    Xcellink Pte Ltd

    Singapore
    28 days ago
  • 4300 - 7100 SGD

     ...Job Summary: The HPC Storage Engineer will be responsible for managing the storage infrastructure within HPC environments. This role involves monitoring storage performance and optimizing through tuning and troubleshooting. Responsibilities: ~ Storage administration... 

    A*STAR RESEARCH ENTITIES

    Singapore
    18 days ago
  • 2800 - 4000 SGD

     ...the need for repairs Obtain comparative quotations Review utilities consumption and strive to minimize costs Supervise foreign staff facilities staff (handy-men) and external contractors Control activities like waste disposal, building security etc. Handle... 

    SSG HOTELS PTE. LTD.

    Singapore
    26 days ago
  • 5000 - 9000 SGD

     ...prototyping (S-Gym), single-GPU workstations (M-Gym), and multi-node HPC at the National Supercomputing Centre Singapore (L-Gym). The...  ...builders around a common platform. We are now hiring full-time engineers and researchers to build, validate, and scale it. Job description... 

    ETH SINGAPORE SEC LTD.

    Singapore
    19 days ago
  • 6000 - 8300 SGD

     ...Position Summary: An Associate End User Computing (EUC) Engineer provides technical support to end users, assisting with hardware and...  ...to have IT procurement experience. Support the procurement for cluster EUC, including submit purchase request, verify invoices, etc.... 

    SYNAPXE PTE. LTD.

    Singapore
    18 days ago
  •  ...Role: Platform Engineer – HPC & Kubernetes Client: Elite FinTech Salary : $120,000 - $250,000 + Bonus Location: Singapore The...  ...GPU resources for intensive computational workloads. Slurm: Proficiency with the Slurm Workload Manager for job scheduling... 

    Hunter Bond

    Singapore
    a month ago
  •  ...work on a next-generation Automotive High Performance Computing (HPC) project, supporting software validation and testing activities to...  ...coverage + Collaborate with software developers, system engineers, and test engineers to ensure software quality and compliance... 

    Aumovio

    Singapore
    6 days ago
  • 5000 - 7000 SGD

     ...AI Infrastructure Engineer 5 days, Mon - Fri 8.30am to 5.30pm Salary...  ...938 Job scopes: Compute & Cluster Management Architect,...  ...orchestration platforms (Kubernetes, Slurm, or Ray) optimized for AI/ML...  ...operator, device plugins) and/or HPC schedulers (Slurm, Run:ai, Ray)... 

    THE SUPREME HR ADVISORY PTE. LTD.

    Singapore
    1 day ago
  •  ...Nscale Nscale is the GPU cloud engineered for AI. We provide cost-...  ...2–3+ years hands-on with GPU, HPC, or large-scale data centre estates...  ...on AI training and inference clusters. Confident with nvidia-smi,...  ...clusters ~ HPC scheduling. Slurm operations for large multi-GPU... 

    Nscale

    Singapore
    more than 2 months ago
  • 9000 - 15000 SGD

     ...replication, and equipment matching. Work with each process cluster to establish critical process windows, statistical control limits...  ...Job Requirements: ~ Bachelor’s or Master’s degree in Engineering (Materials, Chemical, Semiconductor, etc.). ~5+ years of experience... 

    RF360 SINGAPORE PTE. LTD.

    Singapore
    5 days ago
  • 8000 - 9500 SGD

     ...Requirements Minimum Degree in Early Childhood Care & Education and Diploma in Leadership At least 6 years of supervisory and cluster mentorship experience Good communication skills (both oral and written), interpersonal skills with flexibility to work with various... 

    SKOOL4KIDZ PTE. LTD.

    Singapore
    6 days ago
  • 14000 - 23000 SGD

     ...Blockchain to users around the world, through our leading products OKX, OKX Wallet, OKLink and more. What You’ll Be Doing  K8s cluster lifecycle management: Own the build, scaling, version upgrades, daily operations, fault diagnosis, and performance tuning of large-scale... 

    OKBL PTE. LTD.

    Singapore
    18 days ago
  •  ...Job Description & Requirements Support and provide assistance to the Nursing staff in providing the residents with basic hygiene and personal care on a daily basis Assist in Bathing and Changing of Diaper Participates and facilitate in activities with the residents... 

    ABER CARE PTE. LTD.

    Singapore
    10 days ago
  • 2800 - 3800 SGD

     ...team, and behind every successful team is a strong leader. As a Cluster Leader, you will play a key role in shaping patient care,...  ...solving problems and making a meaningful impact on both patients and staff. Key Responsibilities: Lead and Develop People Lead,... 

    HMI ONECARE PTE. LTD.

    Singapore
    20 days ago
  •  ...Opportunity to solve complex performance challenges using C++, HPC About Our Client A leading technology company focused on...  ...edge devices. The team works at the intersection of AI, systems engineering, and hardware optimization. Job Description Develop high-performance... 

    Michael Page

    Singapore
    more than 2 months ago
  • 3000 - 5000 SGD

     ...infrastructure, High Performance Computing (HPC), Enterprise Servers, Storage, Private...  ...are seeking a highly motivated Pre-Sales Engineer who is passionate about technology and enjoys...  ...AI Infrastructure & GPU Computing HPC Clusters & Supercomputing Enterprise Servers & Workstations... 

    NETWEB PTE. LTD.

    Singapore
    more than 2 months ago
  •  ..., including noise, accuracy, bandwidth, power consumption, thermal behavior, and loop stability. Collaborate closely with layout engineers to ensure high-quality analog layout practices, including device matching, parasitic management, and layout-dependent effect mitigation... 

    Skyworks

    Singapore
    1 day ago
  • 1 SGD

     ...MAIN DUTIES: Position Requirements Sales · Work with the Cluster General Manager in implementing the relevant sales strategies and...  ...strict adherence to existing laws, statutes etc. · Ensure all staff within the department work in a manner which is safe and unlikely... 

    THE SINGAPORE RESORT & SPA

    Singapore
    7 days ago
  • 6000 - 8000 SGD

     ...Description To support our extraordinary teams who build great products and contribute to our growth, we’re looking to add a Staff Electrical Engineer located in Kallang, Singapore . This includes maintaining quality engineering programs, standards, and improvements... 

    POWER SYSTEMS R&D (SINGAPORE) PTE. LTD.

    Singapore
    5 days ago
  • 5000 - 9500 SGD

     ...Responsibilities: Manage engineering projects from start to completion, ensuring timely execution and delivery Prepare and maintain project and engineering documentation Review customer specifications and requirements to ensure they meet business and engineering needs... 

    UNITED TEST AND ASSEMBLY CENTER LTD

    Singapore
    8 days ago
  •  ...About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance...  ...escalation point for onsite DC Operations staff; coordinate smart-hands tasks within your...  ...Have High-performance fabrics and GPU-HPC. Exposure to RDMA/InfiniBand, link-level... 

    Nscale

    Singapore
    more than 2 months ago
  • 5400 - 6600 SGD

     ...including Advance Care Planning (ACP), and promote active participation in care decisions. Supervise, mentor, and develop nursing staff through clinical guidance, coaching, orientation, competency assessments, and continuing education. Plan and manage staffing, duty... 

    ABER CARE PTE. LTD.

    Singapore
    25 days ago
  • 16000 - 30000 SGD

     ...integrate, providing a robust billing experience for all users. The engineering team spans across Shanghai and Singapore, offering the chance...  ...and other build tools; 10. Familiar with the use of Docker cluster; 11. Familiar with JSON & XML. Preferred qualifications:... 

    AIRWALLEX (SINGAPORE) PTE. LTD.

    Singapore
    27 days ago
  • 9000 - 11000 SGD

     ...platform teams — defining architecture, aligning technical strategies with business outcomes, and mentoring technical leads and senior engineers to raise the engineering bar. You Will Contribute To Domain and Architectural Leadership Deeply understand core financial... 

    MATCHMOVE PAY PTE. LTD.

    Singapore
    26 days ago
  • 15000 - 28000 SGD

     ...the systems. It is a chance to design from scratch, and create a reliable workflow service at scale. We are looking for a Senior/Staff Engineer to lead the architectural design and implementation of our next-generation durable execution platform. This team is tasked with... 

    COUPANG ASIA HOLDINGS PTE. LTD.

    Singapore
    20 days ago
  • 7000 - 10000 SGD

     ...sustainable, and digitally optimized. Position Overview We are seeking a highly technical, forward-thinking Senior Product Quality Engineer to champion advanced quality planning and reliability engineering. In this role, you will lead the "shift-left" approach to... 

    ENVIRODYNAMICS SOLUTIONS PTE. LTD.

    Singapore
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff Slurm Cluster & HPC Engineer. Be the first to apply!