Sign up to access all features of our service
  • Job search
  • Favorites
  • Create a CV
    New
  • Subscriptions

Linux Infrastructure Engineer

Full-time

Uvation

Job Overview

We are seeking a highly experienced Senior Linux Infrastructure Engineer with deep expertise in Linux administration, bare metal infrastructure, enterprise storage, and next-generation AI Factory / GPU infrastructure platforms . This role is focused on designing, deploying, operating, and troubleshooting large-scale Linux-based infrastructure that powers both traditional enterprise workloads and modern AI/ML environments.

This is not a DevOps-focused role . We already have a dedicated DevOps team and are looking for an engineer with extensive hands-on experience in Bare Metal as a Service (BMaaS), GPU infrastructure, high-performance storage, data center operations, and enterprise Linux platforms .

The ideal candidate will have experience building and managing infrastructure from the hardware layer up, including servers, networking, storage, GPU clusters, and AI-ready platforms. They should be comfortable working with high-performance computing (HPC), AI Factory environments, and large-scale Linux deployments where performance, reliability, and operational excellence are critical.

Key Responsibilities & Required Skills

Linux & Bare Metal Infrastructure

  • Expert-level Linux administration (Ubuntu required; Red Hat and SUSE preferred)
  • Deep expertise in bare metal server deployment, architecture, provisioning, and lifecycle management
  • Experience operating Bare Metal as a Service (BMaaS) platforms and large-scale infrastructure environments
  • Strong understanding of server hardware, including:
    • BIOS/UEFI
    • RAID controllers
    • Firmware management
    • iLO/iDRAC/IPMI
    • NICs and SmartNICs
    • HBA cards
    • Hardware diagnostics and troubleshooting
  • Experience designing, implementing, and supporting enterprise Linux infrastructure at scale

AI Factory & GPU Infrastructure

  • Experience deploying and managing GPU-accelerated infrastructure for AI/ML workloads
  • Understanding of NVIDIA GPU technologies including:
    • A100, H100, H200, B200, or equivalent GPU platforms
    • NVIDIA DGX and OEM GPU servers
    • GPU provisioning and lifecycle management
    • GPU monitoring and performance optimization
  • Knowledge of AI Factory architecture and infrastructure requirements
  • Experience supporting GPU clusters, AI training environments, and high-performance computing (HPC) workloads
  • Understanding of:
    • GPU resource allocation and scheduling
    • Multi-GPU systems
    • GPU networking requirements
    • High-bandwidth, low-latency infrastructure design
  • Familiarity with NVIDIA ecosystem technologies such as:
    • CUDA
    • NCCL
    • GPUDirect Storage
    • NVIDIA Fabric Manager
    • NVIDIA Base Command (preferred)

Enterprise Storage & Data Platforms

  • Advanced Linux storage administration:
    • LVM
    • XFS, EXT4
    • NFS
    • iSCSI
    • Fibre Channel SAN
    • Multipath I/O
  • Strong hands-on experience with Ceph , including:
    • Cluster architecture
    • MON, OSD, MDS
    • RBD, CephFS, RGW
    • Capacity planning
    • Performance tuning
    • Failure recovery
  • Experience with high-performance AI storage platforms such as:
    • WEKA
    • VAST Data
    • Dell PowerScale
    • Pure Storage FlashBlade
    • NetApp
  • Understanding of:
    • NVMe-over-Fabrics (NVMe-oF)
    • RDMA
    • GPUDirect Storage
    • Parallel file systems
    • AI data pipelines

Networking & Infrastructure

  • Strong networking knowledge:
    • Bonding
    • VLANs
    • Routing
    • MTU optimization
    • DNS
    • DHCP
  • Experience with high-performance data center networking:
    • 100G/200G/400G Ethernet
    • RoCE
    • RDMA
    • Spine-Leaf architectures
  • Familiarity with NVIDIA Spectrum-X, Mellanox/NVIDIA ConnectX adapters, or equivalent technologies
  • Strong understanding of Layer 2 and Layer 3 infrastructure design and troubleshooting

Operations & Reliability

  • Experience with high availability, clustering, and disaster recovery
  • Strong troubleshooting skills across:
    • Linux operating systems
    • Hardware platforms
    • GPU infrastructure
    • Networking
    • Enterprise storage
  • Experience supporting mission-critical production environments
  • Bash and Python scripting for automation and operational efficiency
  • Experience creating operational documentation, runbooks, and infrastructure standards
  • Understanding of AI infrastructure design and reference architectures
  • AI cloud integration for workloads
  • SOP and runbook development and maintenance
  • Incident, problem, and capacity management
  • Business continuity and disaster recovery planning for AI workloads
  • Proactive risk identification and mitigation to avoid business impact

Nice to Have

  • Kubernetes infrastructure (especially AI/ML and GPU integration)
  • KVM, VMware, OpenShift Virtualization, or similar virtualization platforms
  • Ansible automation
  • NVIDIA Base Command Manager
  • Slurm or HPC workload schedulers
  • Observability and monitoring platforms (Prometheus, Grafana, OpenTelemetry)
  • Data Center Infrastructure Management (DCIM) tools
  • IPAM solutions
  • AWS, Azure, or hybrid cloud exposure

We Are Not Looking For

  • Candidates whose experience is primarily CI/CD pipeline engineering
  • Engineers focused mainly on Terraform, GitOps, or application delivery pipelines
  • Cloud-only administrators with limited bare metal, storage, or hardware experience
  • Professionals whose primary expertise is software development rather than infrastructure engineering

Ideal Candidate

Someone who has spent years designing, building, and operating enterprise Linux environments, large-scale bare metal infrastructure, storage platforms, and modern AI Factory environments. The ideal candidate understands how to deploy and manage GPU-enabled infrastructure, BMaaS platforms, enterprise storage, and high-performance networking while solving complex operating system, hardware, storage, and AI infrastructure challenges. DevOps experience is a plus, but deep Linux, infrastructure, storage, BMaaS, and AI Factory expertise is the primary requirement.

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Linux Infrastructure Engineer in Remote vacancy
  •  ...Infrastructure Engineer   Location:   Remote – Philippines  Work Schedule:   We're currently hiring for two available shifts:  Dayshift: Monday to Friday, 9 AM to 6 PM PH time  Nightshift: Monday to Friday, 11 PM to 8 AM PH time  About Us   Safari... 
    Remote job

    Safari Micro

    Remote
    25 days ago
  •  ...estimating end to end, from enquiry through to submitting the quote to the client. Remotely, this role owns the estimating and quoting engine room — checking plans, sourcing supplier pricing, working out margins and building the quote — while management keeps final client... 

    Outsourcedstaff

    Remote
    23 days ago
  •  ...We are seeking a skilled and experienced Azure DevOps Engineer with a strong background in Linux administration to join our dynamic team. The ideal...  ...responsible for managing and optimizing our Azure cloud infrastructure, ensuring seamless CI/CD pipelines, and maintaining... 

    Uvation

    Remote
    10 hours ago
  •  ...others. We are looking for a Backend Engineer to join our Core team and help build...  ...of production systems. Contribute to infrastructure improvements, including CI/CD, observability...  .... ~ Experience working with Linux and Docker. ~ Ability to write and maintain... 

    Betterme

    Remote
    10 hours ago
  •  ...are looking for a Full-Stack DevSecOps Engineer to join our globally distributed Engineering...  ...end-user desktop experience to cloud infrastructure, application performance, and security....  ...and optimize our cloud infrastructure, Linux servers, networking, identity management... 
    Remote job

    CQ fluency

    Remote
    25 days ago
  •  ...build what makes people better and keep challenging ourselves to inspire others. Your impact: Manage and optimize cloud infrastructure and resources with attention to cost and resilience. Enhance technical monitoring and alerting to ensure the availability... 

    Betterme

    Remote
    2 days ago
  •  ...We Are Hiring - DevOps / Platform Engineer | 100% Remote   We are seeking a mid-level...  ...production environments rather than building infrastructure from scratch. You'll help improve...  ...supporting production applications (AWS/Azure, Linux, Git, etc.) Experience implementing... 
    Remote job

    Dev Partners

    Remote
    25 days ago
  • Responsibilities Deliver new features and bug fixes with an emphasis on reliability, performance, and user satisfaction. Design and develop responsive, high-performance user interfaces and application state across platforms. Maintain code quality through coding...

    Tether Operations Limited

    Remote
    2 days ago
  •  ...billion in equity and debt funding from global and regional investors, and is now valued at $6,5 billion. We are hiring a Data Engineer for the Data Services Team who will help us create our internal solutions such as a Feature Store and will be ready to perform... 

    Tabby

    Remote
    2 days ago
  • 1000 $ per year

     ...winner (USA & Japan, 2024–2025) and a Top-5 Company for Work-From-Anywhere Jobs (FlexJobs, 2025). We are looking for Senior CV Engineer . Your main tasks will be: Train and fine-tune generative models (text2image, image2image, video2image, IP-adapters) to... 

    Social Discovery Group

    Remote
    10 hours ago
  •  ...data applications work together intelligently. As a Senior Data Engineer, you'll be both a core builder of that platform and a...  ...engineering role. Most of your week is deep technical work — cloud infrastructure, data pipelines, and AI agents. A meaningful part of it is... 
    Remote job

    FinStrat Management

    Remote
    25 days ago
  •  ...impact. We are seeking a Senior Quality Assurance & Test Engineer to join the Engineering team within the Supply Chain and Operations...  ...a high standard of engineering quality, enhancing testing infrastructure and automation processes, and driving continuous improvements... 
    Remote job

    nXscaleSolutions Inc

    Remote
    more than 2 months ago
  •  ...We are seeking a Senior iOS Engineer to own the client side implementation of how members join, pay, and stay with Raya. Member Experience owns the full member lifecycle – applying, onboarding, payments, and lifecycle management – and this role owns the surfaces where... 

    Raya

    Remote
    2 days ago
  • 14400 - 19200 $ per year

     ...Hiring: Full-time Solutions Engineer - Remote - $14,400 - $19,200/yr About the Client Platinum HubSpot Agency Partner – Elite-level HubSpot implementation and revenue architecture expertise. Revenue-Focused Operators – Every solution is engineered to drive... 
    Remote job

    WeAssist.io

    Remote
    25 days ago
  •  ...Remote (Philippines) Work Schedule: Flexible (Supports Global Time Zones) Job Summary We are looking for an experienced AI Engineer to help build the next generation of intelligent applications, AI-powered workflows, and production-ready machine learning solutions... 
    Remote job

    JWay Group

    Remote
    25 days ago
  •  ...Hiring: Full-time - AI Engineer - Remote About us, WeAssist, the hiring team Led by a Founder Who Cares – Reef Colman built WeAssist to empower individuals and create meaningful opportunities that support families. ⚡ Fast & Purposeful Recruiting – We move... 
    Remote job

    WeAssist.io

    Remote
    25 days ago
  •  ...GUIDEWIRE QUALITY ASSURANCE ANALYST / ENGINEER ABOUT THE ROLE We are seeking detail-oriented Guidewire Quality Assurance Analysts or Engineers to support testing and quality assurance across complex insurance transformation engagements. You will play a critical... 
    Remote job

    Norima Consulting

    Remote
    25 days ago
  •  ...an experienced Information and Communications Technology Engineer  for a  long-term independent contractor basis .  This is...  ...systems , including structured cabling, telecommunications infrastructure, security systems, access control, video surveillance, and other... 
    Remote job

    Ehvert Engineering a Salas O'Brien company

    Remote
    25 days ago
  •  ...Senior Network Engineer (Permanent WFH | Night Shift) Location: Remote – Philippines  Shift : 11:00 PM – 8:00 AM PH Time |...  ...technical escalation point for complex network issues, lead infrastructure projects, and partner with cross-functional teams to deliver... 
    Remote job

    Safari Micro

    Remote
    25 days ago
  •  ...experience is an employee first process. Our vision is the same, a place where employees know they can thrive. As a Conversational AI Engineer specializing in Gemini Enterprise for Customer Experience (GECX), you will design, deploy, and optimize next-generation, customer-... 

    Ttecdigital

    Remote
    29 days ago
  •  ...of driven, independent self-learners who are pushing the boundaries of Digital Construction. If you are are an experienced BIM engineer with the ability to problem-solve inventively, develop new workflows and enjoy collaborating in small focused teams, then DigitalMatter... 
    Remote job

    DigitalMatter

    Remote
    25 days ago
  •  ...help businesses worldwide improve operations and achieve growth. Position: Mid & Senior Oracle Utilities QA, Test & Performance Engineer The QA & Performance Engineer will be responsible for end-to-end quality assurance, test scenario creation, test execution,... 
    Remote job

    BRight Co., Ltd.

    Remote
    25 days ago
  •  ...About the Front-End Developer & Software Engineer position We currently run a SaaS application that has large businesses as clients and need a front-end developer / software engineer who is capable of migrating software from AngularJS across to a single page application... 
    Remote job

    B2B HQ

    Remote
    25 days ago
  •  ...responsible for advanced troubleshooting, systems administration, infrastructure support, cybersecurity remediation, and helping drive...  ...multiple priorities. Team Collaboration Work closely with engineers, project teams, account managers, and security personnel.... 
    Remote job

    Nexplay Consulting Inc.

    Remote
    25 days ago
  •  ...progress through completion Coordinate with Support, NOC, Engineering, and other technical teams to implement security...  ...identifying opportunities to strengthen workstation, cloud, infrastructure, and security configurations Support system hardening initiatives... 
    Remote job

    Atlas Technica

    Remote
    25 days ago
  •  ...Reporting directly to the Head of Technology and the CEO , this role sits at the intersection of strategic planning, hands-on infrastructure engineering, and daily internal client satisfaction. You will oversee the complete lifecycle of our networks, hardware systems, and... 
    Remote job

    Skyrocket Studios PH, Inc

    Remote
    more than 2 months ago
  •  ...that power Skyrocket Studios’ digital products and marketing infrastructure. Where You’ll Lead This role combines advanced website...  ...to development best practices. Cloud, AI & Data Engineering Develop and integrate AI-powered features and agents using... 
    Remote job

    Skyrocket Studios PH, Inc

    Remote
    more than 2 months ago
  •  ..., SQL/PL-SQL, HQL, COBOL-toJava migrations, and Oracle Cloud Infrastructure (OCI) environment architecture. ~ Proven track record acting...  ...clients, managing offshore teams of 5–15+ developers and engineers. ~ Certified Oracle Utilities Implementation Professional (... 
    Remote job

    BRight Co., Ltd.

    Remote
    25 days ago
  •  ...will be responsible for developing and maintaining automation infrastructure that connects multiple business tools, reduces manual work,...  ...generation workflows Optimize AI outputs using prompt engineering and iterative testing Technical Implementation & Integration... 
    Remote job

    VA Desk

    Remote
    25 days ago
  •  ...Position: Senior Full Stack Engineer (Java/React)  Location: Remote from LATAM  Contract Type: Full-time vendor  Time Zone...  ...focuses on the modernization and scaling of critical backend infrastructure, specifically regarding high-volume transaction handling and... 
    Remote job

    In All Media Inc

    Remote
    25 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Linux Infrastructure Engineer. Be the first to apply!