Lead Staff Systems Reliability Engineer (Linux & Distributed Systems)

The Trade Desk

London, UKonsitePosted Oct 2, 2025
Posting intelligenceMay be filled, listed long ago

Skills

kubernetesprometheusmongonextgenansiblegopythonkafkacssrustrubyc#

About the role

The Trade Desk is a global technology company and the world’s leading independent platform for digital advertising, with nearly 4,000 employees across more than 30 offices. Our technology helps advertisers reach the right audiences across the open internet — from streaming TV and podcasts to mobile apps, news, and more.

Advertising powers the content people love. By making it more transparent, effective, and responsible, we help support trusted journalism, quality entertainment, and creators worldwide. The world’s brands and agencies rely on us to reach their customers and grow their businesses responsibly.

The scale of our platform brings unique technical challenges — from processing massive datasets in real time to building systems that operate reliably on a global scale. When you work here, your impact is worldwide. We welcome diverse perspectives, encourage curiosity, and build teams that learn from one another. If you’re driven to solve meaningful challenges, we’d love to meet you.

What we do

We are looking to hire a Lead Systems Reliability Engineer to join our engineering team to continue building and maintaining our data-driven platform. We leverage technologies like Aerospike, MongoDB, and Kafka to perform many real time activities, translating to with a p99 latency under 1 millisecond on the back end!

Do you enjoy tuning, performance testing, troubleshooting, automation, and operating at scale? Does testing next-gen hardware, evaluating data access patterns, and designing automation around distributed systems excite you?

What makes this role different:

First in the Industry: The Trade Desk is the first company to run over 5MM QPS to NVMe in Aerospike on a single node, forcing core software redesigns to achieve this scale.

Work on Cutting-Edge Hardware: Design clusters with nodes featuring 300TB of NVMe, 3TB RAM, and 512 cores, delivering a global 2,500GB/s throughput directly from flash.

Shape the Future of Infrastructure: Spec your own systems and collaborate directly with AMD and NoSQL vendors to run PoCs and optimize bleeding-edge technology for internet-scale workloads.

Deep Performance Engineering: Dive into kernel, hardware, and system interactions, leveraging tools like flamegraphs, NUMA counters, BIOS tuning, and synthetic testing to achieve world-class performance.

Push Hardware Endurance Limits: Build clusters engineered to withstand over 1 zettabyte of endurance.

What you’ll do:

Lead a team to influence, manage, and plan work streams, systems, and data structures at scale within a global ecosystem, spanning multiple infrastructure providers (cloud and traditional datacenters).

Encourage, improve, and build infrastructure automation in a way that works with stateful systems at scale.

Own operations for Linux-based systems running Aerospike, Kafka, and Mongo.

Serve as a point of contact to review new use cases, answer questions, and participate in on-call rotation.

Learn to be a NoSQL SME. You do not need experience to apply – we will train you.

Benchmark and analyze next generation hardware offerings.

Who you are:

Skills and Experience

Linux operating system

Leadership experience and ability to mentor

Troubleshooting

Techniques for isolation, scientific method

Identify bottlenecks (Is it CPU? IO?)

Nice-To-Have experience:

Physical hardware (on-prem) internals, management, and operation

Performing testing and tuning

Databases (relational or NoSQL)

Ansible/PyInfra/Chef

Prometheus

Kubernetes

Python/Ruby/Rust/Bash/Golang/C#

An Empathetic, Objective, Critical Thinker:

Thinking beyond the task at hand to deeply understand the 'why' behind an objective.

A welcoming of ideas, and understanding of, perspectives that are different from your own and an interest in seeking and building from a common ground.

You are a creative thinker, not bound by "the way things have always been done" but are thinking of the questions nobody has thought of and are "yet to be asked". What you know is less important than how well you learn, innovate, collaborate, and adapt.

As a global team from many diverse backgrounds, experiences, and perspectives, you value and seek out paths for fostering diversity.

Please reach out to us at accommodations@thetradedesk.com to request an accommodation or discuss any accessibility needs you may require to access our Company Website or navigate any part of the hiring process.

When you contact us, please include your preferred contact details and specify the nature of your accommodation request or questions. Any information you share will be handled confidentially and will not impact our hiring decisions.

Questions about this role

Click "Apply with AI Applyd" above. We auto-fill the application from your resume and answer screening questions in seconds. No copy and paste, no juggling tabs.

Compensation for Software Engineer roles in United Kingdom varies widely by seniority, employer size, and remote vs onsite arrangement. Check the salary range on this listing when published, or browse our Software Engineer hub for United Kingdom medians across recent openings.

Most applications complete in under 90 seconds. You can track the status in your dashboard and watch the screenshot proof land the moment the application submits.

AI Applyd supports Greenhouse, Lever, Ashby, Workday, iCIMS, SmartRecruiters, Personio, Teamtailor and other major ATS platforms. If we can submit through the platform, we do.

Want AI Applyd to auto-apply to roles like this?

We tailor your resume per posting, fill the forms, and track replies for you.