Site Reliability Engineer, Hybrid Cloud Operation and Delivery - Data Infrastructure
12000 - 22000 SGDBYTEDANCE PTE. LTD.
About Us
Founded in 2012, ByteDance's mission is to inspire creativity and enrich life. With a suite of more than a dozen products, including TikTok, Lemon8, CapCut and Pico as well as platforms specific to the China market, including Toutiao, Douyin, and Xigua, ByteDance has made it easier and more fun for people to connect with, consume, and create content.
Why Join ByteDance
Inspiring creativity is at the core of ByteDance's mission. Our innovative products are built to help people authentically express themselves, discover and connect – and our global, diverse teams make that possible. Together, we create value for our communities, inspire creativity and enrich life - a mission we work towards every day.
As ByteDancers, we strive to do great things with great people. We lead with curiosity, humility, and a desire to make impact in a rapidly growing tech company. By constantly iterating and fostering an "Always Day 1" mindset, we achieve meaningful breakthroughs for ourselves, our Company, and our users. When we create and grow together, the possibilities are limitless. Join us.
Diversity & Inclusion
ByteDance is committed to creating an inclusive space where employees are valued for their skills, experiences, and unique perspectives. Our platform connects people from across the globe and so does our workplace. At ByteDance, our mission is to inspire creativity and enrich life. To achieve that goal, we are committed to celebrating our diverse voices and to creating an environment that reflects the many communities we reach. We are passionate about this and hope you are too.
Responsibilities
Team Introduction
Our team is responsible for infrastructure systems of hybrid cloud, including products in IaaS/PaaS/SaaS/AI models. We strive to be a leading Site Reliability Engineering (SRE) team in the industry, driving reliability, scalability, and performance at scale.As part of the SRE team, you will tackle complex, large-scale challenges, leveraging your expertise in coding, algorithms, complexity analysis, and distributed system design.We foster a culture of diversity, intellectual curiosity, and open collaboration. Engineers are empowered with strong ownership, autonomy, and the opportunity to work across a wide range of impactful projects.
What you will be doing:
1. Responsible for delivery products in hybrid cloud scenarios, including cloud platform planning, software deployment, resource expansion, etc. Collaborate with R&D teams to complete project delivery.
2. Responsible for the operation of cloud platform environments for internal and external customers, including daily alarm handle, on-call support, change, as well as ensuring stability of cloud platform during important event periods.
3. Participate in stability construction of cloud products with R&D team, and continuously improve capabilities in high availability architecture, disaster recovery, alarm monitoring, etc, based on the experience we get from large-scale systems on site.
4. Continuously promote the improvement of hybrid cloud serviceability, participate in the standardized SOW of O&M and delivery for new product versions, and build the SRE serviceability acceptance standards to improve implement efficiency.
Qualifications
Minimum Qualifications:
- Bachelor's / Master's Degree in Computer Science or related major, with at least 5 years of relevant experience;
- Solid basic knowledge of computer software, understanding of Linux operating system, network ,middleware and other related principles.
- Familiar with one or more programming languages, such as Shell, Python, Go, or Java. Knowledge of building scripts or tools to handle different problems.
- Experience in operation and maintenance of one or more fields, including virtual machines, containers, K8s, load balancing, middleware, AI models, etc.
Preferred Qualifications
- Experience in operation and maintenance of IDC equipments such as switches and GPU servers
- Working experience in cloud platform related vendors
8000 - 12000 SGD
...Design, implement, operate, and maintain secure, reliable, and resilient infrastructure across the Company’s data centres, corporate offices, and public cloud environments. Administer and support enterprise compute infrastructure, including Linux and Windows servers, virtualization...6000 - 7000 SGD
...Raffles Place Mon to Fri 9am-6pm Your Responsibilities: - Design, implement, operate, and maintain secure, resilient infrastructure across data centres, corporate environments, and public cloud platforms. Administer Windows/Linux servers, compute, virtualization,...5000 - 11000 SGD
...Platform Operations Engineer (Hybrid Infrastructure) Role Overview The Platform Operations Engineer supports on-premises and hybrid infrastructure platforms that underpin mission-critical systems. This role focuses on platform reliability, operational stability, and...8000 - 9500 SGD
...and maintain observability solutions across applications, infrastructure, cloud and hybrid environments. Work with metrics, logs, events and... ...Identify monitoring gaps and recommend improvements to system reliability and performance. Work with engineering, infrastructure...- ...youth-at-risk, underprivileged families, and elderly, to conserving the environment. Job Purpose The Site Reliability Engineer (SRE) drives enterprise operational resilience by architecting scalable cloud infrastructure, managing the enterprise DevOps toolchain, and...
14000 - 28000 SGD
...Responsibilities: Infrastructure Automation & Configuration Management... ...manage virtual machines using cloud-init, QCOW2 images, and... ...premises, including control plane operations, node provisioning, workload... ...AWS ECR) Cloud Platforms Engineer and maintain infrastructure...7000 - 9000 SGD
...We are seeking an experienced Site Reliability Engineer (SRE) / Application Support Engineer to support and maintain critical banking applications... ...incidents, driving service improvements, and enhancing operational efficiency through automation and SRE best practices. Key...3500 - 5500 SGD
...and supply chain. We are currently looking to hire SRE Engineer (Site Reliability Engineer). This is an exciting opportunity to expand your... ...troubleshoot telemetry issues, integrate observability into delivery pipelines, and drive adoption of standard metrics, logs,...7000 - 9500 SGD
...conversion) About the Role We are looking for an experienced Site Reliability Engineer (SRE) to support and enhance a large-scale government... ...ensuring platform reliability, scalability, security, and operational efficiency through automation, monitoring, and continuous improvement...8000 - 16000 SGD
...We are looking for seasoned engineers to join our team! You will participate... ...implementation, deployment, operational management, and debugging.... .... Design and deliver infrastructure solutions that will help build... ...responsibility and ownership of F5 Cloud Platform software components,...20000 - 40000 SGD
...people who join us now. The opportunity The SVP/ED, Site Reliability Engineering is a senior strategic executive leadership role responsible... ...’s enterprise reliability, observability, automation, and operational resilience agenda across critical platforms and services....8000 - 10000 SGD
...disruptions with a structured, data-driven approach to minimise... ...and resolve application and infrastructure issues, provide timely... ...recurrence. Maintain the reliability, availability, and performance... ...reduce manual effort, improve operational resilience, perform routine...5500 - 10000 SGD
...Production Support / Site Reliability Engineer (SRE) Role Overview We are looking for an experienced Production Support / Site Reliability... ...is a hands-on, techno-functional role covering production operations, incident resolution, system reliability, and stakeholder...10000 - 15000 SGD
...Responsibilities 1.Responsible for the operations, monitoring, and resource management of big data solutions on overseas cloud platforms, ensuring the stability and reliability of data platforms and related services. 2.Develop and improve automated operations tools to...12000 - 20000 SGD
...and trading firm that is expanding its engineering footprint in Tokyo, Japan. The team sits... ...heart of a systematic trading business operating across equities, options, and other... ...experience in: Production Engineering / Site Reliability Engineering Electronic Trading...10000 - 15000 SGD
...support large-scale digital capabilities and high-volume data operations. about the role As an Engineering Manager, you will lead a team of data engineers... .... You will set technical standards, manage core infrastructure, and align data architecture with business goals....9000 - 10500 SGD
...We are looking for a Cloud Infrastructure Engineer to design, implement, administer and support enterprise cloud and hybrid infrastructure environments, primarily using Microsoft Azure . Key Responsibilities Design, deploy and manage Azure IaaS/PaaS infrastructure...16000 - 20000 SGD
...looking for an experienced and hands-on, Vice President of Site Reliability Engineering to lead the reliability, availability, and resilience strategy... ..., and error budgets , ensuring reliability decisions are data-driven. Drive improvements in service availability,...4500 - 9000 SGD
...We're looking for a Cloud Infrastructure Engineer to design, implement, and support secure and resilient cloud infrastructure powering critical government... ...Why Join: Contribute to Government Digital Projects Hybrid Work Arrangements Up to S$9,000 + AWS What you'll do:...7000 - 9500 SGD
...Location: Punggol (Hybrid) Job Type: Full-Time Contract (Subjected to renewal/conversion) Join a high-impact engineering team supporting large-scale digital government platforms. As a Site Reliability Engineer (SRE), you will help ensure platform reliability, scalability...19000 - 21000 SGD
...Overview We are seeking an experienced Delivery Leader to drive successful delivery and adoption of IT Infrastructure Services (Cloud), ensuring customer satisfaction, retention... ...adoption of new technologies such as AIOps, hybrid cloud, and DevOps. Ensure service excellence...7000 - 10000 SGD
...Summary We are seeking an experienced Production Support Engineer / Site Reliability Engineer (SRE) to maintain availability and stability of a... ...financial services environment, driving automation and operational resilience. Responsibilities Maintain availability and...8000 - 16000 SGD
...world's most complex and mission-critical systems. As a Site Reliability Engineer III at JPMorgan Chase within the Payments Technology team,... ...with simple and straightforward solutions. Through code and cloud infrastructure, you will configure, maintain, monitor, and optimize...9000 - 14000 SGD
...AWS Network Engineer – Cloud Infrastructure & Security Role Overview We are seeking a hands-on AWS Network Engineer to design, implement, secure, and maintain cloud and hybrid network infrastructure supporting enterprise applications and mission-critical systems....7000 - 9000 SGD
...key IT industry drivers such as security, cloud, data management and IoT, can address customer... ...such as revenue growth and business, operational efficiency, innovation, risk and compliance... ...Job summary: The Cloud Operations Engineer is an Operations role who will focus on...5500 - 9500 SGD
...Overview We're hiring a Site Reliability Engineer to join a high-performing AI Platform team responsible for building and operating the core infrastructure, runtime and operational foundations that enable AI adoption across the organisation. This role is ideal for someone...4000 - 6000 SGD
...Job description We are looking for a Cloud & Infrastructure Engineer to build, deploy and secure the technology infrastructure behind our Industrial... ...and project teams to turn applications into secure, reliable and production-ready systems. Roles & Responsibilities...7000 - 8000 SGD
...Site Reliability Engineer (SRE) Location: Singapore, Onsite Employment Type: Contract (12 months renewable based project demand / performance... ..., and SLOs for critical production services. Automate operational processes, enhance runbooks, and minimise manual support effort...6500 - 12000 SGD
...systems, ensuring compliance with privacy and data protection laws, managing changes and releases, and driving operational automation and infrastructure optimization. High Availability and... .... # At least 5 years of experience in cloud services products and application...- ...is the AI-native financial operating system for a real-time, intelligent... ...in 2015 to build the infrastructure global commerce runs on. We'... ...About the team The Engineering team at Airwallex is a diverse... ...together to build scalable, reliable, and secure products that empower...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer, Hybrid Cloud Operation and Delivery - Data Infrastructure. Be the first to apply!
