hirly

Apply with hirly

Engineering Team Lead, Site Reliability Engineering (Digital Banking)

Mtb · Wilmington, DE

Upload your resume to see how well you match this job — free, in seconds, no account needed.

Your resume is used only to score it against this job. If you don't create an account, it is deleted within 24 hours.

Join M&T Bank's Digital Banking organization and play a critical leadership role in ensuring the reliability, resiliency, and performance of the platforms our customers depend on every day. As Engineering Team Lead, Site Reliability Engineering (SRE), you will lead a team of engineers responsible for the stability and operational excellence of M&T's online and mobile banking platforms, supporting millions of customer interactions across digital channels. This role is ideal for a technology leader who thrives in a highly visible, customer-impacting environment and is passionate about building highly available, resilient systems. You will drive Site Reliability Engineering best practices, observability strategies, incident management processes, service-level performance, and release engineering practices while helping the organization modernize its digital banking ecosystem. Working across both legacy and modern platforms, you will partner with engineering, architecture, infrastructure, and business teams to improve reliability, scalability, and customer experience. The successful candidate will combine strong technical expertise in cloud-native and distributed systems, Microsoft Azure, Infrastructure as Code (IaC), CI/CD automation, and modern release engineering practices with proven leadership capabilities, helping shape the future of digital banking through operational excellence, automation, resiliency engineering, and continuous improvement. Primary Responsibilities

  • Lead a team of Site Reliability Engineers responsible for the availability, resiliency, performance, and operational health of M&T's Digital Banking platforms.
  • Drive implementation and adoption of Site Reliability Engineering best practices, including service level objectives (SLOs), service level agreements (SLAs), error budgets, observability, monitoring, and automation.
  • Partner with engineering teams to design, build, and support highly available and fault-tolerant applications and services.
  • Establish and mature incident management, problem management, root cause analysis, and operational readiness processes.
  • Develop and execute strategies that improve platform stability, reliability, scalability, and customer experience across mobile and online banking channels.
  • Lead reliability initiatives supporting both legacy platforms and modern technology stacks as Digital Banking continues its transformation journey.
  • Oversee release readiness, production support activities, release governance, and operational governance for customer-facing applications.
  • Partner with business and technology stakeholders to prioritize reliability investments and align operational objectives with business goals.
  • Monitor application health, performance metrics, and customer-impacting incidents while driving continuous service improvements.
  • Serve as a subject matter expert for observability, monitoring, reliability engineering, release engineering, and operational excellence.
  • Drive adoption of cloud engineering and reliability best practices across Microsoft Azure environments.
  • Partner with engineering teams to implement and maintain Infrastructure as Code (IaC) solutions utilizing Terraform and automated platform provisioning practices.
  • Lead continuous improvement of software delivery processes through GitLab, Artifactory, CI/CD automation, and deployment pipeline optimization.
  • Establish and oversee release engineering practices, including release readiness, deployment governance, rollback procedures, and production validation processes.
  • Design and promote resilient deployment strategies including Active-Active architectures, Blue/Green deployments, Canary releases, and other progressive delivery methodologies.
  • Guide and mentor engineers while fostering a culture of accountability, innovation, collaboration, and continuous learning.
  • Manage staffing plans, workload allocation, team development, and resource planning across multiple initiatives.
  • Evaluate emerging technologies, vendor solutions, and industry trends to improve operational efficiency and engineering effectiveness.
  • Ensure adherence to enterprise architecture standards, technology governance processes, and operational risk requirements.
  • Manage project priorities, delivery timelines, and technology investments within assigned areas.
  • Build strong partnerships across Digital Banking, Engineering, Infrastructure, Architecture, Risk, and Product organizations.
  • Exercise usual authority of a manager concerning staffing, performance appraisals, promotions, salary recommendations, performance management, and terminations.
  • Understand and adhere to the Company's risk and regulatory standards, policies and controls in accordance with the Company's Risk Appetite. Design, implement, maintain and enhance internal controls to mitigate risk on an ongoing basis. Identify risk-related issues needing escalation to management.
  • Promote an environment that supports belonging and reflects the M&T Bank brand.
  • Maintain M&T internal control standards, including timely implementation of internal and external audit points together with any issues raised by external regulators as applicable.
  • Complete other related duties as assigned. Scope of Responsibilities Oversees a team where the majority of employees are Site Reliability Engineers, Software Engineers, and technical individual contributors supporting Digital Banking platforms. Responsible for the reliability, observability, operational excellence, release management, deployment governance, and supportability of customer-facing online and mobile banking applications. Supervisory/Managerial Responsibilities
  • 5 to 10 direct reports Education and Experience Required
  • A combined minimum of 9 years' higher education and/or work experience, including a minimum of 4 years' engineering and/or architecture experience and 3 years leadership experience
  • Capable of working on multiple projects of a complex nature
  • Proficiency with project management, word processing and spreadsheet applications
  • Complete understanding of the system development life cycle
  • Excellent problem-solving skills to assist in issue resolution
  • Familiar with application development software and hardware platforms
  • Excellent verbal and written communication skills
  • Excellent analytical skills
  • Excellent decision-making skills
  • Strong project management skills
  • Strong presentation skills
  • Experience encouraging teamwork and serving as role model when leading and directing others
  • Understanding of technical, business and operational impacts of a project or problem
  • Experience with Site Reliability Engineering principles, operational excellence, and production support practices
  • Experience leading technical teams supporting highly available customer-facing applications
  • Experience with application monitoring, observability, logging, alerting, and incident management processes
  • Knowledge of service reliability metrics including SLOs, SLAs, availability, performance, and resiliency objectives
  • Experience supporting distributed systems, cloud-based platforms, or modern application architectures
  • Experience supporting Microsoft Azure environments and cloud-native technologies
  • Experience with Infrastructure as Code (IaC) tools such as Terraform
  • Experience implementing and supporting CI/CD pipelines and deployment automation practices
  • Experience with source control, artifact management, and software delivery tooling such as GitLab and Artifactory
  • Knowledge of release engineering, deployment automation, and production release governance practices
  • Understanding of high-availability deployment models, including Active-Active architecture and zero-downtime deployment strategies
  • Experience collaborating with cross-functional technology and business stakeholders Education and Experience Preferred
  • Bachelor's degree
  • Minimum of 10 years' technology management or large program leadership exp
Apply: Engineering Team Lead, Site Reliability Engineering (Digital Banking) at Mtb