Sign up to access all features of our service
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Linux Infrastructure Engineer

Full-time

Uvation

Job Overview

We are seeking a highly experienced Senior Linux Infrastructure Engineer with deep expertise in Linux administration, bare metal infrastructure, enterprise storage, and next-generation AI Factory / GPU infrastructure platforms . This role is focused on designing, deploying, operating, and troubleshooting large-scale Linux-based infrastructure that powers both traditional enterprise workloads and modern AI/ML environments.

This is not a DevOps-focused role . We already have a dedicated DevOps team and are looking for an engineer with extensive hands-on experience in Bare Metal as a Service (BMaaS), GPU infrastructure, high-performance storage, data center operations, and enterprise Linux platforms .

The ideal candidate will have experience building and managing infrastructure from the hardware layer up, including servers, networking, storage, GPU clusters, and AI-ready platforms. They should be comfortable working with high-performance computing (HPC), AI Factory environments, and large-scale Linux deployments where performance, reliability, and operational excellence are critical.

Key Responsibilities & Required Skills

Linux & Bare Metal Infrastructure

  • Expert-level Linux administration (Ubuntu required; Red Hat and SUSE preferred)
  • Deep expertise in bare metal server deployment, architecture, provisioning, and lifecycle management
  • Experience operating Bare Metal as a Service (BMaaS) platforms and large-scale infrastructure environments
  • Strong understanding of server hardware, including:
    • BIOS/UEFI
    • RAID controllers
    • Firmware management
    • iLO/iDRAC/IPMI
    • NICs and SmartNICs
    • HBA cards
    • Hardware diagnostics and troubleshooting
  • Experience designing, implementing, and supporting enterprise Linux infrastructure at scale

AI Factory & GPU Infrastructure

  • Experience deploying and managing GPU-accelerated infrastructure for AI/ML workloads
  • Understanding of NVIDIA GPU technologies including:
    • A100, H100, H200, B200, or equivalent GPU platforms
    • NVIDIA DGX and OEM GPU servers
    • GPU provisioning and lifecycle management
    • GPU monitoring and performance optimization
  • Knowledge of AI Factory architecture and infrastructure requirements
  • Experience supporting GPU clusters, AI training environments, and high-performance computing (HPC) workloads
  • Understanding of:
    • GPU resource allocation and scheduling
    • Multi-GPU systems
    • GPU networking requirements
    • High-bandwidth, low-latency infrastructure design
  • Familiarity with NVIDIA ecosystem technologies such as:
    • CUDA
    • NCCL
    • GPUDirect Storage
    • NVIDIA Fabric Manager
    • NVIDIA Base Command (preferred)

Enterprise Storage & Data Platforms

  • Advanced Linux storage administration:
    • LVM
    • XFS, EXT4
    • NFS
    • iSCSI
    • Fibre Channel SAN
    • Multipath I/O
  • Strong hands-on experience with Ceph , including:
    • Cluster architecture
    • MON, OSD, MDS
    • RBD, CephFS, RGW
    • Capacity planning
    • Performance tuning
    • Failure recovery
  • Experience with high-performance AI storage platforms such as:
    • WEKA
    • VAST Data
    • Dell PowerScale
    • Pure Storage FlashBlade
    • NetApp
  • Understanding of:
    • NVMe-over-Fabrics (NVMe-oF)
    • RDMA
    • GPUDirect Storage
    • Parallel file systems
    • AI data pipelines

Networking & Infrastructure

  • Strong networking knowledge:
    • Bonding
    • VLANs
    • Routing
    • MTU optimization
    • DNS
    • DHCP
  • Experience with high-performance data center networking:
    • 100G/200G/400G Ethernet
    • RoCE
    • RDMA
    • Spine-Leaf architectures
  • Familiarity with NVIDIA Spectrum-X, Mellanox/NVIDIA ConnectX adapters, or equivalent technologies
  • Strong understanding of Layer 2 and Layer 3 infrastructure design and troubleshooting

Operations & Reliability

  • Experience with high availability, clustering, and disaster recovery
  • Strong troubleshooting skills across:
    • Linux operating systems
    • Hardware platforms
    • GPU infrastructure
    • Networking
    • Enterprise storage
  • Experience supporting mission-critical production environments
  • Bash and Python scripting for automation and operational efficiency
  • Experience creating operational documentation, runbooks, and infrastructure standards
  • Understanding of AI infrastructure design and reference architectures
  • AI cloud integration for workloads
  • SOP and runbook development and maintenance
  • Incident, problem, and capacity management
  • Business continuity and disaster recovery planning for AI workloads
  • Proactive risk identification and mitigation to avoid business impact

Nice to Have

  • Kubernetes infrastructure (especially AI/ML and GPU integration)
  • KVM, VMware, OpenShift Virtualization, or similar virtualization platforms
  • Ansible automation
  • NVIDIA Base Command Manager
  • Slurm or HPC workload schedulers
  • Observability and monitoring platforms (Prometheus, Grafana, OpenTelemetry)
  • Data Center Infrastructure Management (DCIM) tools
  • IPAM solutions
  • AWS, Azure, or hybrid cloud exposure

We Are Not Looking For

  • Candidates whose experience is primarily CI/CD pipeline engineering
  • Engineers focused mainly on Terraform, GitOps, or application delivery pipelines
  • Cloud-only administrators with limited bare metal, storage, or hardware experience
  • Professionals whose primary expertise is software development rather than infrastructure engineering

Ideal Candidate

Someone who has spent years designing, building, and operating enterprise Linux environments, large-scale bare metal infrastructure, storage platforms, and modern AI Factory environments. The ideal candidate understands how to deploy and manage GPU-enabled infrastructure, BMaaS platforms, enterprise storage, and high-performance networking while solving complex operating system, hardware, storage, and AI infrastructure challenges. DevOps experience is a plus, but deep Linux, infrastructure, storage, BMaaS, and AI Factory expertise is the primary requirement.

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Linux Infrastructure Engineer in Home based vacancy
  •  ...Job Overview  :  We are looking for a highly experienced Linux Infrastructure Engineer with deep expertise in traditional Linux administration, bare metal infrastructure, and enterprise storage. This role is not focused on DevOps or cloud-native engineering. We already... 
    Remote job

    Uvation

    Home based
    25 days ago
  •  ...Join the company that’s building the telemetry infrastructure for the AI era. At Cribl, we partner with IT and Security teams at many of the...  ...Location: Singapore Why You'll Love This Role As a Solutions Engineer, you'll get to flex your technical chops and storytelling skills... 

    Cribl

    Home based
    5 days ago
  •  ...building one of the most modern and globally accessible financial infrastructure platforms in the industry, built to advance an open, global...  ...clients. The Growth Foundations team builds the AI-native Growth Engine, enabling Growth teams across the company. We design and ship... 

    Kraken

    Home based
    1 day ago
  •  ...What You'll Do We are seeking a talented Senior FullStack Engineer to join our engineering team and take ownership of building our...  ...schemas using PostgreSQL Deploy and manage applications on AWS infrastructure with focus on scalability and performance Collaborate with... 
    Remote job

    CITA LABS

    Home based
    25 days ago
  •  ...for clients worldwide. What You'll Do As a Senior Java Engineer at CITA Lab, you will play a key role in designing and developing...  ..., payment systems, digital asset platforms, and blockchain infrastructure. You will work on systems where performance, reliability, security... 
    Remote job

    CITA LABS

    Home based
    25 days ago
  •  ...Job Overview We are looking for an experienced Network Security Engineer to design, implement, monitor, and support enterprise security infrastructure across on-premises, cloud, and hybrid environments. The ideal candidate should possess strong expertise in next-generation... 
    Remote job

    Uvation

    Home based
    25 days ago
  •  ...applications. - Design highly available and fault-tolerant database infrastructure suitable for production-critical systems. - Optimize complex...  ..., and performance analysis. - Work closely with backend engineers on transaction boundaries, locking strategies, concurrency... 
    Remote job

    CITA LABS

    Home based
    25 days ago
  •  ...building one of the most modern and globally accessible financial infrastructure platforms in the industry, built to advance an open, global...  ..., and payment providers. You’ll be joining the Solutions Engineering function as a core contributor on our pre-sales team. Your focus... 

    Kraken

    Home based
    1 day ago
  •  ...and Android capabilities, portable libraries and shared mobile infrastructure. The ambition is simple to describe and difficult to execute:...  ...a single framework. We’re also interested in Principal-level engineers who have solved comparable problems through Flutter, Kotlin Multiplatform... 

    Canva

    Home based
    4 days ago
  •  ...production issues across applications, databases, networks, and infrastructure. - Participate in architecture reviews and technical design discussions. - Conduct code reviews and establish backend engineering standards. - Write comprehensive unit, integration, and end-... 
    Remote job

    CITA LABS

    Home based
    25 days ago
  •  ...You'll Do We are looking for a Senior Rust & Smart Contract Engineer to design, build, and maintain secure, high-performance...  ...systems and smart contracts. You will work on core blockchain infrastructure, on-chain protocols, and integrations that interact with decentralized... 
    Remote job

    CITA LABS

    Home based
    25 days ago
  •  ...Intelligence, Application Development, Smart City Technology, Digital Infrastructure, and Cybersecurity. At GovTech, we offer you a purposeful...  ...standards with ease. What you will be working on : AI Engineers in the Responsible AI team will focus on executing safety... 
    Remote job

    National Library Board

    Home based
    more than 2 months ago
  •  ...powerful ways to monitor, troubleshoot, and optimize their AI systems. That’s where we come in. Arize AI is the leading AI & Agent Engineering observability and evaluation platform , empowering AI engineers to ship high-performing, reliable agents and applications. From first... 

    Arizeai

    Home based
    3 days ago
  •  ...and system state during testing. - Identify, document, prioritize, and track defects. - Reproduce production issues and work with engineers to identify root causes. - Design test scenarios covering edge cases, failure conditions, and abnormal user behavior. - Perform... 
    Remote job

    CITA LABS

    Home based
    25 days ago
  •  ...and graceful degradation. - Ensure compatibility across modern browsers and different screen sizes. - Work closely with backend engineers to define API contracts and integration patterns. - Collaborate with UI/UX designers to improve usability and product consistency.... 
    Remote job

    CITA LABS

    Home based
    25 days ago
  •  ...ABOUT US At LangChain, our mission is to make intelligent agents ubiquitous. We build the foundation for agent engineering in the real world, helping developers move from prototypes to production-ready AI agents that teams can rely on. We began as widely adopted open-source... 

    Langchain

    Home based
    11 days ago
  •  ...detect sophisticated attacks, investigate risk quickly, and automate response at enterprise scale. We're building a team of strong engineers who are excited to use modern AI development tools to build faster, make better decisions, and ship high-quality software with more... 

    Abnormalsecurity

    Home based
    3 days ago
  •  ...distributed systems implemented in Go and running on Kubernetes across all major Cloud service providers (GCP, Azure, AWS). We also have engineers working on Prometheus , Grafana Alloy , Mimir proxies , and OpenTelemetry. As a company we are remote-first and global, we embrace... 

    Grafanalabs

    Home based
    3 days ago
  •  ...Join the company that’s building the telemetry infrastructure for the AI era. At Cribl, we partner with IT and Security teams at many of the...  ...our value proposition up and down the organization, from engineer up to CxO - Forecasting predictably and hitting sales targets... 

    Cribl

    Home based
    5 days ago
  •  ...is the trusted source for open source. By delivering hardened, secure, and production-ready builds of all the open source software engineers and AI agents rely on, Chainguard helps organizations build faster, stay compliant, and eliminate risk. Our customers include Fortune... 

    Chainguard

    Home based
    3 days ago
  •  ...initial requirements through design, development, testing, release, and continuous improvement. You will work closely with clients, engineers, designers, QA, DevOps, Security, and other stakeholders to define what we build, why we build it, and how it should work.... 
    Remote job

    CITA LABS

    Home based
    25 days ago
  • 1000 SGD per annum

     ...notice and remotely. Support and maintain an AlwaysOn availability group (2 servers: primary + mirror). Help analysts and ML engineers with data access and retrieval. Run routine database maintenance: archiving, recovery strategies, security policies, data integrity... 

    Social Discovery Group

    Home based
    1 day ago
  •  ...worldwide. We are looking for a motivated and solution-oriented engineer who enjoys solving technical challenges and driving continuous...  ...and software development, is highly preferred. ~ Proficiency in Linux shell scripting/Python/SQL, and AI tools for automation and data... 

    Mellanox Technologies

    Home based
    28 days ago
  •  ...experienced Batch Operations & Automation Engineer to support Business-as-Usual (BAU)...  ...will work closely with application and infrastructure teams to design and implement batch solutions...  ...understanding of Windows, Unix, and Linux operating systems . Understanding of... 

    Quesscorp Singapore Pte Ltd

    Home based
    a month ago
  •  ...Position: Salesforce Engineer Location: Remote from APAC Contract Type: Full-time Time Zone Alignment: APAC About In All Media: We are a Managed Nearshore Teams provider headquartered in Austin, specializing in building and embedding high-performing software... 
    Remote job

    In All Media Inc

    Home based
    25 days ago
  •  ...not afraid to pick up the phone Nice to have: Bachelor's degree in a relevant field or experience as a radiographer, field service engineer or similar. We offer a competitive salary and benefits package, as well as opportunities for career growth and development. If you... 
    Remote job

    AlemHealth

    Home based
    25 days ago
  •  ...invests in. If you are deeply fluent in insurance product suites and platform roadmaps, comfortable holding both a head of claims and an engineer in the same conversation, and driven to keep clients and internal teams honest to the outcomes they signed up for - we'd love to hear... 

    Sapiens

    Home based
    a month ago
  •  ...initiatives in close collaboration with Regional Sales, Product Management, R&D, and Technology teams to secure technical specifications at engineering lead levels. 4. Translate deep Voice of Customer insights into concrete NPI proposals and application-specific technology... 

    Morgan Philips Group

    Home based
    a month ago
  •  ...Job Description  Microsoft Copilot Studio Power Platform Engineer   Mandatory Skills:  Copilot studio development, custom agent creation, powerplatform, powerapps, power automate, azure AIML experience, python and devops   Role Overview   We are seeking an experienced... 

    Kasi Jeyaseelan Naveen (Proprietor of Arient Solutions)

    Home based
    2 days ago
  •  ...NVIDIA is looking for a Senior Network Engineer to join the Global Backbone Engineering team. Powering NVIDIA's DGX Cloud - a distributed...  ...platform - the team operates the network enabling the infrastructure's worldwide presence. They architect and scale the long-haul,... 

    Mellanox Technologies

    Home based
    28 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Linux Infrastructure Engineer. Be the first to apply!