AI Engineer (ML Systems & Infrastructure)
16000 - 25000 SGDSWAPETECH PTE. LTD.
About the Role
We are looking for exceptional AI Engineers to build the next generation of AI infrastructure and Machine Learning Systems(MLSys).
This role focuses on large-scale system infrastructure rather than model research. You will work on the core foundations that power large-scale AI training and inference systems, including Kubernetes cluster management, RDMA networking, unified KV Cache architecture, observability platforms, distributed systems, GPU orchestration, and CUDA kernel optimisation.
You will collaborate closely with AI researchers, infrastructure architects, networking engineers, and platform teams to maximize the efficiency, scalability, and reliability of AI systems.
Key Responsibilities
AI Infrastructure & Kubernetes
- Design, deploy, and operate large-scale Kubernetes-based AI infrastructure.
- Develop cluster governance frameworks, scheduling policies, resource isolation, and multi-tenancy capabilities.
- Build and optimize GPU orchestration platforms using Kubernetes, Slurm, Volcano, Kueue, Ray, and related technologies.
- Improve cluster utilization, reliability, elasticity, and operational efficiency.
RDMA & High-Performance Networking
- Design and optimize RDMA, InfiniBand, RoCE, and high-speed Ethernet fabrics for distributed AI workloads.
- Optimize GPU-to-GPU and GPU-to-NIC communication paths.
- Improve distributed communication efficiency for large-scale training and inference.
- Analyze and eliminate networking bottlenecks across AI clusters.
Unified KV Cache & Distributed Memory Systems
- Design and implement unified KV Cache architecture across:
- GPU HBM
- CPU Memory
- RDMA-accessible Memory
- NVMe SSD
- Distributed Storage
- Develop efficient KV Cache sharing, migration, offloading, and scheduling mechanisms.
- Optimize latency and throughput for large-scale inference systems.
CUDA & System Performance Optimisation
- Develop and optimize CUDA kernels for training and inference workloads.
- Profile and optimize GPU compute, memory, communication, and scheduling efficiency.
- Contribute to low-level optimization of AI frameworks and inference engines.
- Work on technologies such as FlashAttention, TensorRT, Triton, NCCL, CUTLASS, and custom operators.
Observability & Reliability
- Build end-to-end observability platforms for AI infrastructure.
- Design monitoring, logging, tracing, alerting, and troubleshooting frameworks.
- Develop performance dashboards and SLO-driven operational systems.
- Improve maintainability, debuggability, and operational excellence of AI platforms.
Automation & Platform Engineering
- Build automation tools for deployment, provisioning, monitoring, and operations.
- Develop Infrastructure-as-Code (IaC) solutions using Terraform, Ansible, and related tools.
- Build CI/CD pipelines and engineering productivity platforms.
- Improve platform scalability and operational efficiency.
Required Qualifications
Education
- Bachelor's degree or above in Computer Science, Software Engineering, Electrical Engineering, or related fields.
Technical Skills
- Strong software engineering and programming skills.
- Excellent system design capability and strong engineering craftsmanship.
- Strong coding standards and code quality awareness.
- Strong sense of ownership, accountability, and execution.
System Fundamentals
Strong understanding of:
- Operating Systems
- Computer Networks
- Distributed Systems
- Data Structures and Algorithms
- Linux Internals
Programming Languages
Proficiency in one or more of:
- C++
- Go
- Python
- Rust
AI Infrastructure Experience
Hands-on experience in one or more of:
- Kubernetes
- GPU Infrastructure
- Distributed Systems
- AI Infrastructure
- HPC (High Performance Computing)
- Cloud-Native Platforms
Networking Experience
Experience with:
- RDMA
- InfiniBand
- RoCE/RoCEv2
- GPUDirect
- NCCL
- UCX
- High-Speed Ethernet
GPU & Performance Engineering
Experience with:
- CUDA
- GPU Performance Optimization
- Multi-GPU Systems
- Distributed Training
- Distributed Inference
Preferred Qualifications
- Experience building large-scale AI training or inference clusters.
- Experience with vLLM, SGLang, TensorRT-LLM, Triton, DeepSpeed, Megatron-LM, Ray, or similar frameworks.
- Experience with unified KV Cache systems, memory hierarchy optimisation, or distributed storage systems.
- Experience with Kubernetes GPU Operator and NVIDIA NetworkOperator.
- Experience with Prometheus, Grafana, Loki, OpenTelemetry, and observability platforms.
- Experience contributing to open-source projects such as: vLLM, FlashAttention, CUTLASS, TVM, MLIR, Triton, Kubernetes, NCCL
- Experience working across AI Infrastructure, HPC, Networking, and Silicon Systems is highly desirable.
- ...groundbreaking research and development in AI and Data Science. Founded in... ...for sensible and inquisitive engineers and scientists with a strong... ...high-performance computing infrastructure. Apply your domain... ...Experience with parallel computing, system programming. Benefits...
5000 - 7000 SGD
...As a Software Engineer (ML/AI), You will work closely with the team to build and enhance systems that support proactive identification of fraudulent behavior across our clientele’s platforms. This is an opportunity to be part of a high-impact team that combines data, AI, and...- ...The AI/ML Engineer / Senior AI/ML Engineer is a highly technical, hands-on role at the intersection of Enterprise AI and client-facing architecture... ...take into account critical enterprise requirements—including systems-level hardening, infrastructure scalability, and network...
6000 - 8500 SGD
...Overview We are seeking a Applied AI Engineer to embed directly with our business units... ...deliverable solutions Build agentic AI systems • Design and build production-grade agentic... ...AI workloads Maintain & enhance ML/DL models • Own, maintain, and improve production...5000 - 6500 SGD
...TacnIQ.ai is bridging AI and the physical world , building the foundation model for physical... .... We're creating a simple platform for engineers to build physical AI with any sensor, across... ...for someone at the beginning of their AI/ML career who wants to work on real, novel problems...- ...Work at the forefront of Agentic AI, and large-scale machine learning system exposure to model training, GPU infrastructure, and AI platform engineering About Our Client Our client is a global technology leader that develops advanced computing, data, and intelligent...
5000 - 9000 SGD
...Project background The AIS Instrumentation Gym is a multi-tier ML platform being built by both the National University... ...scientists, ML researchers, and systems builders around a common platform. We are now hiring full-time engineers and researchers to build, validate, and...6000 - 12000 SGD
...Data Engineer (AI Infrastructure) About the Role Our client is a fast-growing AI technology company looking for a Data Engineer to build and optimize... ...(AWS, Azure, GCP). Experience with GPU computing or AI/ML infrastructure. Familiarity with Docker, Kubernetes, or CI...6000 - 8000 SGD
...help turn model workflows into practical product features. You will work closely with the founding team, product engineering, and production users to build AI systems that support real creative decision-making. What you'll do Build and evaluate AI workflows for script-to...5800 - 7000 SGD
...Responsible for installation, configuration, testing and support in IT system implementation projects. Manage IT operations including... ..., active directory, backups, virtualization, storage and other infrastructure related tools. Understanding the customers system...- ...Software Engineer (Distributed Systems - Python) Role: Software Engineer (Distributed Systems - Python) Client: Elite Tech Firm Compensation... ...high-performance, distributed systems for large-scale ML infrastructure. Key responsibilities include: Design and build...
7000 - 14000 SGD
...will suit an experienced Application/IT system administrator with experience of building... ...and performing upgrades of system infrastructure. The role will involve the design, installation... ...qualification in Computer Science or Engineering or equivalent experience [update qualification...5500 - 6500 SGD
...(eligible for CAT1 security clearance). Education: Diploma or Degree in Information Technology, Electrical/Electronic Engineering, Information Systems, or a related discipline. Experience: Minimum 3 years of hands-on experience in systems administration, system configuration...5000 - 7000 SGD
...lifecycle of Machine Learning and AI projects, delivering data-... ...definition, data exploration, feature engineering, model training, validation,... ...Develop and implement ML, AI, or optimization models to... ...Manage AI projects and maintain AI infrastructure to ensure operational...- ...Job Summary We are looking for an experienced AI Engineer with strong expertise in Generative AI , AI/ML , or Microsoft Copilot/Agentic AI development. Candidates... ...Vision, Predictive Analytics, or Recommendation Systems. Model training, evaluation, deployment, and...
4000 - 4500 SGD
...-growing digital agency building AI-powered products. Our AI team is... ...are looking for a skilled DevOps Engineer to help us scale, automate, and secure our infrastructure. Role Overview You will be... ...cloud environments that power our AI/ML workflows. This role is ideal for...5000 - 9000 SGD
...Project background The AIS Instrumentation Gym is a multi-tier ML platform being built by both the National University... ...scientists, ML researchers, and systems builders around a common platform. We are now hiring full-time engineers and researchers to build, validate, and...6400 - 9500 SGD
...spatial perception products for autonomous and tele-operated robotic systems, focusing on lightweight sensor hardware, advanced vision-based... .... We are looking for curious, mathematically strong engineers or scientists who enjoy solving difficult problems and translating...4000 - 7000 SGD
...Support the deployment and implementation of AI infrastructure and data centre projects. Coordinate... ..., technology partners, and internal engineering teams. Assist in the design and... ...solutions, including GPU servers, storage systems, networking equipment, and cooling...8000 - 9200 SGD
...Job Summary The Manager, AI / MI Ops is responsible for managing the day-to-day operations... ...applications, dashboards, and reporting systems operate reliably, meet business... ...reporting accuracy by collaborating with Data Engineering and business stakeholders. Manage user...5500 - 8000 SGD
...We are looking for an experienced HPC Systems Engineer to support and operate large-scale Linux... ..., and maintain Linux-based HPC infrastructure, including compute nodes, storage platforms... ...workloads Support compute-intensive, AI, and data-driven applications Advise...3500 SGD
...About the role We are seeking a motivated AI/ML Engineer to joinour growing AI and software development team. Working closely with our CTO... ...and maintainmachine learning models, LLM-based and agentic-AI systems, and AI-poweredbusiness applications that drive our digital transformation...- ...connected devices and embedded systems across all industries,... ...manufacturing facilities, critical infrastructure, and government entities. We... ...Development (AD) Reports To: Engineering Manager / Tech Lead... ...are looking for a mid-level AI/ML Software Engineer to join...
3400 - 3900 SGD
...System Engineer (Network & Infrastructure) Salary Range: $3,400 - $3,900 Working Days: 5 days, Monday - Friday Working Hours: 8.30am to 5.30pm Location: Jurong East (Toh Guan) Responsibilities Identify, research, and develop suitable technologies to enhance...- ...applied intelligence that brings AI into the real world. Our... ...simulation, and modern software engineering to accelerate the design,... ...the data and machine-learning infrastructure the platform runs on: the pipelines... ...models can train on, and the systems that move those models from...
8000 - 10000 SGD
...The role Gengis AI is building applied AI for the operational layer... .... You will own the core systems, AI orchestration, evaluation workflows, and product infrastructure that make this possible. You will... ...narrow backend role and not a pure ML research role; you should be able...7000 - 12000 SGD
...AIOps & Agentic AI Engineering Develop Python and/or Lang Graph scripts and agentic AI workflows... ...intelligent event correlation across infrastructure and applications. Design and... ...auto-remediation workflows driven by AI/ML and event correlation insights. Infrastructure...5000 - 6000 SGD
...Core Experience: Must have hands-on experience in System Administration for production systems , with a strong focus on Linux... ...(Oracle, SQL Server, MySQL) and JDBC drivers. Modern Infrastructure : Experience with Cloud platforms , Docker , Kubernetes...6500 - 9500 SGD
...Design, automate and manage Azure cloud infrastructure using Infrastructure as Code (Terraform)... ...GitOps and modern cloud-native platform engineering practices. Develop monitoring, logging... ...and operational visibility. Support AI initiatives by enabling AI platforms, integrating...10000 - 12000 SGD
...Role Overview The AI/ML Specialist will play a pivotal role in designing and implementing... ...speech recognition, NLP, recommendation systems, and time series forecasting. Lead the... ...strategies for AI solutions in cloud infrastructure at scale. Drive technical design...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Engineer (ML Systems & Infrastructure). Be the first to apply!
