AI Engineer (ML Systems & Infrastructure)
16000 - 25000 SGDSWAPETECH PTE. LTD.
About the Role
We are looking for exceptional AI Engineers to build the next generation of AI infrastructure and Machine Learning Systems(MLSys).
This role focuses on large-scale system infrastructure rather than model research. You will work on the core foundations that power large-scale AI training and inference systems, including Kubernetes cluster management, RDMA networking, unified KV Cache architecture, observability platforms, distributed systems, GPU orchestration, and CUDA kernel optimisation.
You will collaborate closely with AI researchers, infrastructure architects, networking engineers, and platform teams to maximize the efficiency, scalability, and reliability of AI systems.
Key Responsibilities
AI Infrastructure & Kubernetes
- Design, deploy, and operate large-scale Kubernetes-based AI infrastructure.
- Develop cluster governance frameworks, scheduling policies, resource isolation, and multi-tenancy capabilities.
- Build and optimize GPU orchestration platforms using Kubernetes, Slurm, Volcano, Kueue, Ray, and related technologies.
- Improve cluster utilization, reliability, elasticity, and operational efficiency.
RDMA & High-Performance Networking
- Design and optimize RDMA, InfiniBand, RoCE, and high-speed Ethernet fabrics for distributed AI workloads.
- Optimize GPU-to-GPU and GPU-to-NIC communication paths.
- Improve distributed communication efficiency for large-scale training and inference.
- Analyze and eliminate networking bottlenecks across AI clusters.
Unified KV Cache & Distributed Memory Systems
- Design and implement unified KV Cache architecture across:
- GPU HBM
- CPU Memory
- RDMA-accessible Memory
- NVMe SSD
- Distributed Storage
- Develop efficient KV Cache sharing, migration, offloading, and scheduling mechanisms.
- Optimize latency and throughput for large-scale inference systems.
CUDA & System Performance Optimisation
- Develop and optimize CUDA kernels for training and inference workloads.
- Profile and optimize GPU compute, memory, communication, and scheduling efficiency.
- Contribute to low-level optimization of AI frameworks and inference engines.
- Work on technologies such as FlashAttention, TensorRT, Triton, NCCL, CUTLASS, and custom operators.
Observability & Reliability
- Build end-to-end observability platforms for AI infrastructure.
- Design monitoring, logging, tracing, alerting, and troubleshooting frameworks.
- Develop performance dashboards and SLO-driven operational systems.
- Improve maintainability, debuggability, and operational excellence of AI platforms.
Automation & Platform Engineering
- Build automation tools for deployment, provisioning, monitoring, and operations.
- Develop Infrastructure-as-Code (IaC) solutions using Terraform, Ansible, and related tools.
- Build CI/CD pipelines and engineering productivity platforms.
- Improve platform scalability and operational efficiency.
Required Qualifications
Education
- Bachelor's degree or above in Computer Science, Software Engineering, Electrical Engineering, or related fields.
Technical Skills
- Strong software engineering and programming skills.
- Excellent system design capability and strong engineering craftsmanship.
- Strong coding standards and code quality awareness.
- Strong sense of ownership, accountability, and execution.
System Fundamentals
Strong understanding of:
- Operating Systems
- Computer Networks
- Distributed Systems
- Data Structures and Algorithms
- Linux Internals
Programming Languages
Proficiency in one or more of:
- C++
- Go
- Python
- Rust
AI Infrastructure Experience
Hands-on experience in one or more of:
- Kubernetes
- GPU Infrastructure
- Distributed Systems
- AI Infrastructure
- HPC (High Performance Computing)
- Cloud-Native Platforms
Networking Experience
Experience with:
- RDMA
- InfiniBand
- RoCE/RoCEv2
- GPUDirect
- NCCL
- UCX
- High-Speed Ethernet
GPU & Performance Engineering
Experience with:
- CUDA
- GPU Performance Optimization
- Multi-GPU Systems
- Distributed Training
- Distributed Inference
Preferred Qualifications
- Experience building large-scale AI training or inference clusters.
- Experience with vLLM, SGLang, TensorRT-LLM, Triton, DeepSpeed, Megatron-LM, Ray, or similar frameworks.
- Experience with unified KV Cache systems, memory hierarchy optimisation, or distributed storage systems.
- Experience with Kubernetes GPU Operator and NVIDIA NetworkOperator.
- Experience with Prometheus, Grafana, Loki, OpenTelemetry, and observability platforms.
- Experience contributing to open-source projects such as: vLLM, FlashAttention, CUTLASS, TVM, MLIR, Triton, Kubernetes, NCCL
- Experience working across AI Infrastructure, HPC, Networking, and Silicon Systems is highly desirable.
6500 - 11500 SGD
...Our team builds the foundational infrastructure that empowers Machine Learning Engineers to develop the next generation of... ...the high-performance, large-scale systems that process petabytes of vehicle... ...and deploy core components of our ML infrastructure platform on Kubernetes...13000 - 16000 SGD
...Role We are seeking a skilled Machine Learning Engineer to join our AI team and build cutting-edge recommendation systems for gaming content. You will leverage state-of-... ...translate business requirements into practical AI/ML solutions that enhance overall user experience....5700 - 6400 SGD
...Kubernetes, Slurm, or Ray) optimized for AI/ML distributed workloads. Monitor GPU health... ...and scale high-throughput parallel file systems and object storage (e.g., Lustre, GPFS/... ...high-speed data pipelines. 3. Automation &Infrastructure as Code (IaC) Build and manage...4500 - 5500 SGD
...passionate and highly motivated Fresh Graduates to join our AI & Machine Learning Engineering team. This role offers an exciting opportunity to work on... ...Semantic Search and intelligent knowledge retrieval systems. Develop and optimize Retrieval-Augmented Generation (RAG...6000 - 8500 SGD
...Role Overview We are seeking a Applied AI Engineer to embed directly with our business units... ...solutions Build agentic AI systems • Design and build production-grade agentic... ...support AI workloads Maintain & enhance ML/DL models • Own, maintain, and improve...11000 - 18000 SGD
...implementation of Red Hat OpenShift AI platform architecture.... ...junior consultants and customer engineers. Engage in pre-sales activities... ...management. ~ Experience with AI/ML frameworks and libraries such... ..., data sources, and business systems. ~ Experience developing secure...- ...Work on cutting-edge AI infrastructure and heterogeneous GPU system High-impact role shaping next-generation large-scale AI and LLM system About Our... ...workloads worldwide. The organisation is known for its strong engineering culture, innovation in AI platforms, and commitment to...
6000 - 10000 SGD
...Responsibilities · Support design of new system infrastructure for client projects · Install, configure and maintain new / existing systems... .../maintenance Requirements · Degree Computer Science or Engineering or equivalent experience · Excellent knowledge of...6000 - 8400 SGD
...IT Infrastructure Engineer Job Description Implement projects to ensure service delivery covering... ...IDM, etc) - Manage platform operations system architecture for new and existing applications... ...applications. - Experience with AI and machine learning tools, platforms, with...8000 - 12000 SGD
...Machine Learning Engineer (AI Infrastructure) Join an AI startup building an intelligent productivity assistant that helps users automate tasks and... ...using advanced AI. Key responsibilities Build and maintain ML/AI infrastructure for training, evaluation, deployment, and...7500 - 10500 SGD
...establishing a brand-new, strategic AI department. This department is... ...Service (MaaS) product support system. Based on massive multi-country... ...and CI/CD of the core product infrastructure Requirements: Master’s in Computer science, Engineering, Information systems, or a...5500 - 8000 SGD
...We are looking for an experienced HPC Systems Engineer to support and operate large-scale Linux... ..., and maintain Linux-based HPC infrastructure, including compute nodes, storage platforms... ...workloads Support compute-intensive, AI, and data-driven applications Advise...- ...Software Engineer (Distributed Systems - Python) Role: Software Engineer (Distributed Systems - Python) Client: Elite Tech Firm Compensation... ...high-performance, distributed systems for large-scale ML infrastructure. Key responsibilities include: Design and build...
- ...groundbreaking research and development in AI and Data Science. Founded in... ...for sensible and inquisitive engineers and scientists with a strong... ...high-performance computing infrastructure. Apply your domain... ...Experience with parallel computing, system programming. Benefits...
3500 - 4800 SGD
...About the Role: We are looking for a hands-on and motivated IT Systems & Infrastructure Engineer to join our MIS team in Singapore. This is a broad infrastructure role covering Microsoft365, identity and access management, endpoint security, servers, backup and recovery...6000 - 18000 SGD
...Our client is a high-growth AI company AI solutions helping enterprises... ...are looking for a Senior AI Engineer to drive the design and delivery of production-grade AI systems. Job Responsibilities Design... ...components Integrate AI / ML models into robust, real-world...5000 - 9000 SGD
...Project background The AIS Instrumentation Gym is a multi-tier ML platform being built by both the National University... ...scientists, ML researchers, and systems builders around a common platform. We are now hiring full-time engineers and researchers to build, validate, and...- ...Work at the forefront of Agentic AI, and large-scale machine learning system exposure to model training, GPU infrastructure, and AI platform engineering About Our Client Our client is a global technology leader that develops advanced computing, data, and intelligent...
10000 - 13000 SGD
...Role Overview The AI/ML Specialist will play a pivotal role in designing and implementing... ...speech recognition, NLP, recommendation systems, and time series forecasting. Lead the... ...strategies for AI solutions in cloud infrastructure at scale. Drive technical design reviews...6000 - 12000 SGD
...Data Engineer (AI Infrastructure) About the Role Our client is a fast-growing AI technology company looking for a Data Engineer to build and optimize... ...(AWS, Azure, GCP). Experience with GPU computing or AI/ML infrastructure. Familiarity with Docker, Kubernetes, or CI...8000 - 10000 SGD
...The client is an innovative AI Fintech company on a mission to... ...experiences, and recommendation engines through backend services and APIs... ...skills. · Comfortable building systems that operate under latency... ...behavior. · Experience with ML Kit, TensorFlow Lite, or lightweight...3400 - 3900 SGD
...System Engineer (Network & Infrastructure) Salary Range: $3,400 - $3,900 Working Days: 5 days, Monday - Friday Working Hours: 8.30am to 5.30pm Location: Jurong East (Toh Guan) Responsibilities Identify, research, and develop suitable technologies to enhance...6000 - 7500 SGD
...environments. • Ensure high availability, redundancy, and system performance across infrastructure components. • Manage VMware/Hyper‑V platforms and... ...tasks. • Closely collaborate with Project Manager and Engineers to support the fulfilment of contract obligations within...8000 - 12500 SGD
...# Domain Knowledge: # Mandatory experience in AML and Compliance domains Familiarity with data analytics and tools such as Actimize, Detica, Fircosoft, Quantexa, Tookitaki Desirable Skills: Exposure to AI/ML technologies and workflow applications is a plus....4000 - 6000 SGD
...About the Role We're building AI-powered solutions that run directly on... ...wearable devices and vehicle-based systems among them. This is a hands-on engineering role for someone who's comfortable at... ...• Demonstrated experience deploying ML/AI models on resource-constrained or...8000 - 13000 SGD
...Our client is a fast-growing AI infrastructure company building next-generation GPU-powered AI platforms. They operate high-performance AI... ...closely with Platform, Infrastructure, Linux, DevOps, and AI Engineering teams to deliver scalable AI infrastructure. Develop and maintain...4000 - 6500 SGD
...Manage and support enterprise Windows and Linux server infrastructure to ensure system availability and stability. Perform server deployment,... ...incidents. Requirements Diploma/Degree in IT, Computer Engineering, Electronics Engineering, or a related field. Minimum...5500 - 8500 SGD
...handover materials. Monitor model in production and maintain the model performance. Follow privacy, security, data-retention and AI-governance requirements. Minimum Required Experience and Qualifications Practical Python, deep-learning and Computer Vision experience...- ...design and build of mission-critical, real-time detection systems Own and shape cutting-edge AI/ML-driven detection strategies About Our Client Our... ...workflows for critical use cases. Partner across engineering, data, product, risk, and operations; lead and grow a...
5000 - 10000 SGD
...Summary We are looking for an ML Evaluation Engineer/Data Scientist to bring statistical rigor and measurement excellence to the evaluation of AI and LLM systems. This role will focus on ensuring that evaluation results are reliable, reproducible, and scientifically valid...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Engineer (ML Systems & Infrastructure). Be the first to apply!
