Site Reliability Engineer
Susquehanna International Group Dublin, IrelandSite Reliability Engineer
Susquehanna International Group Dublin, Ireland
Overview
We're looking for a Site Reliability Engineer to keep the trading platforms we depend on healthy, performant and continuously improving. You'll operate within a set of platforms that span Windows and Linux: a distributed platform built on Service Fabric, real-time news infrastructure, and a large Python production environment, running across development, integration, and production in multiple regions. When something breaks you'll dive in, diagnose quickly, and fix it - then make the change that stops it happening again.
This is a hands-on engineering role that sits at the boundary between development and infrastructure. The team is evolving into reliability engineering: automation-first, monitoring-as-code, self-healing systems, and shared tooling built to production standard. You'll work in a tight-knit, business-facing team alongside developers, production and infrastructure engineers in Dublin, and partner closely with global platform-engineering teams - consuming their products and contributing to them upstream rather than building regional one-offs.
In this role you will:
OPERATE Help run the production estate day to day: deployments, platform operations, capacity, certificates, and health checks across environments and tenants. Triage and resolve incidents quickly, perform root-cause analysis, and drive the permanent fix so the same problem doesn't come back.
AUTOMATE Take an automation-first approach to everything the team touches: build auto-remediation for known failure modes, replace manual runbooks with code, and treat recurring toil as a defect to be engineered away.
OBSERVE Operate the monitoring and alerting stack (Checkmk, Elasticsearch/Kibana, Grafana, Prometheus): move its configuration into code, tune signal from noise, and make sure every alert that fires is actionable.
RELEASE Own safe delivery to production: design and maintain CI/CD pipelines in GitLab and deployment orchestration in Octopus Deploy.
ENGINEER Build and productionise the team's tooling with CI, tests, and documentation as standard, and contribute improvements upstream to the global platform products the estate runs on rather than maintaining local one-offs.
COLLABORATE Act as the bridge between development teams and the infrastructure teams (compute, network, storage, data-centre operations), applying your knowledge of both worlds: translate what developers need into infrastructure terms and infrastructure constraints back into design decisions, and build tooling that can be handed off to infrastructure teams to run.
What we're looking for
Susquehanna is an equal opportunity employer. If you need assistance and/or a reasonable accommodation due to a disability during the application or recruitment process, please contact us.
If you're a recruiting agency and want to partner with us, please reach out to recruiting@sig.com . Any resume or referral submitted in the absence of a signed agreement will not be eligible for an agency fee.
We're looking for a Site Reliability Engineer to keep the trading platforms we depend on healthy, performant and continuously improving. You'll operate within a set of platforms that span Windows and Linux: a distributed platform built on Service Fabric, real-time news infrastructure, and a large Python production environment, running across development, integration, and production in multiple regions. When something breaks you'll dive in, diagnose quickly, and fix it - then make the change that stops it happening again.
This is a hands-on engineering role that sits at the boundary between development and infrastructure. The team is evolving into reliability engineering: automation-first, monitoring-as-code, self-healing systems, and shared tooling built to production standard. You'll work in a tight-knit, business-facing team alongside developers, production and infrastructure engineers in Dublin, and partner closely with global platform-engineering teams - consuming their products and contributing to them upstream rather than building regional one-offs.
In this role you will:
OPERATE Help run the production estate day to day: deployments, platform operations, capacity, certificates, and health checks across environments and tenants. Triage and resolve incidents quickly, perform root-cause analysis, and drive the permanent fix so the same problem doesn't come back.
AUTOMATE Take an automation-first approach to everything the team touches: build auto-remediation for known failure modes, replace manual runbooks with code, and treat recurring toil as a defect to be engineered away.
OBSERVE Operate the monitoring and alerting stack (Checkmk, Elasticsearch/Kibana, Grafana, Prometheus): move its configuration into code, tune signal from noise, and make sure every alert that fires is actionable.
RELEASE Own safe delivery to production: design and maintain CI/CD pipelines in GitLab and deployment orchestration in Octopus Deploy.
ENGINEER Build and productionise the team's tooling with CI, tests, and documentation as standard, and contribute improvements upstream to the global platform products the estate runs on rather than maintaining local one-offs.
COLLABORATE Act as the bridge between development teams and the infrastructure teams (compute, network, storage, data-centre operations), applying your knowledge of both worlds: translate what developers need into infrastructure terms and infrastructure constraints back into design decisions, and build tooling that can be handed off to infrastructure teams to run.
What we're looking for
- Proven experience in a site reliability, DevOps, platform, or production engineering role, operating systems a business depends on.
- Strong scripting and automation skills (e.g. Python, PowerShell, Bash).
- Comfortable working in both Windows and Linux environments.
- Willingness to work shifts: the team works a rotating early/late weekday pattern, with occasional weekend work some participation in an on-call rotas.
- Hands-on experience with CI/CD tooling (e.g. GitLab) and deployment orchestration platforms (e.g. Octopus Deploy).
- Experience with configuration-management and automation frameworks (e.g. Ansible).
- Experience building and tuning monitoring and alerting (e.g. Checkmk, Elasticsearch/Kibana, Grafana, Prometheus).
- Strong troubleshooting instincts and root-cause discipline - not satisfied until the class of problem is fixed, not just the instance.
- Strong communication skills - able to convey technically complex subjects clearly to developers and partner teams.
- Understanding and practical application of security models (e.g. RBAC) and certificate management.
- A working understanding of the infrastructure layer - compute, networking, storage, identity - and how software behaves on top of it.
- B.S. in a technical discipline or equivalent experience.
Susquehanna is an equal opportunity employer. If you need assistance and/or a reasonable accommodation due to a disability during the application or recruitment process, please contact us.
If you're a recruiting agency and want to partner with us, please reach out to recruiting@sig.com . Any resume or referral submitted in the absence of a signed agreement will not be eligible for an agency fee.
Job ID 10903
More Jobs From Susquehanna International Group
Susquehanna International Group
Dublin, Ireland
Susquehanna International Group
Dublin, Ireland
Susquehanna International Group
Dublin, Ireland
Susquehanna International Group
Dublin, Ireland
Susquehanna International Group
Dublin, Ireland
Susquehanna International Group
Sydney, Australia
Susquehanna International Group
Bala-Cynwyd, United States
Susquehanna International Group
Bala-Cynwyd, United States
Susquehanna International Group
Bala-Cynwyd, United States
Susquehanna International Group
Bala-Cynwyd, United States
Boost your career
Find thousands of job opportunities by signing up to eFinancialCareers today.More Jobs Like This
Broadridge Trading & Connectivity Solutions
Manila, Philippines