
Senior Application Support Engineer (SRE)
Depository Trust & Clearing Corporation (DTCC)11 days ago
Tampa, FL, USASenior
Responsibilities
- Partner with engineering and infrastructure teams to improve application reliability, resilience, observability, and operational excellence.
- Lead critical production incident resolution, impact analysis, root-cause investigations, preventive actions, and post-incident reviews.
- Own incident, problem, change, and major incident management processes.
- Identify reliability risks and implement solutions that prevent recurrence and reduce operational toil.
- Develop and maintain runbooks, knowledge articles, and operational documentation.
- Support production releases, deployment activities, vendor upgrades, disaster recovery testing, execution, and audit readiness.
- Drive automation and alert optimization initiatives to improve efficiency and reduce operational noise.
- Apply risk, control, security, and reliability best practices across day-to-day operations.
- Collaborate with global technology, infrastructure, operations, and business teams to maintain high availability.
Requirements
- At least 6 years of experience in application support, Site Reliability Engineering, or production engineering.
- Bachelor’s degree preferred or equivalent experience.
- Hands-on application production support experience focused on reliability, observability, production stability, resiliency, and incident prevention.
- Extensive experience supporting Linux and Windows environments, including process inspection, log analysis, troubleshooting, performance diagnostics, and system optimization.
- Strong scripting and automation experience with Bash, Shell Scripting, Python, Ruby, Perl, and JavaScript.
- Experience with enterprise monitoring, observability, analytics, and testing platforms including Dynatrace, Splunk, Grafana, and Selenium.
- Technical expertise with SQL databases such as Oracle, Snowflake, PostgreSQL, or similar technologies, including data analysis, production troubleshooting, root-cause investigation, and query optimization.
- Experience with ServiceNow and Jira and with Incident, Problem, Change, and Major Incident Management processes.
- Experience supporting enterprise messaging and queueing technologies such as IBM MQ, Oracle AQ, ActiveMQ, RabbitMQ, and Kafka.
- Experience with enterprise job scheduling platforms, particularly Autosys.
- Working knowledge of OpenShift, AWS, Amazon RDS Aurora, PostgreSQL, and cloud-native operational support practices.
- Familiarity with COBOL, JCL, DB2, DB2 Stored Procedures, CICS, SPUFI, and File-AID.
- Understanding of certificate management, password management, access controls, security, risk, and operational controls.
- Capital markets or financial services experience supporting mission-critical, highly available, and regulated systems is expected.
- Knowledge of artificial intelligence concepts and their practical application in production support, operational engineering, and service reliability.
- Strong communication, stakeholder management, leadership, ownership, analytical, problem-solving, and critical-thinking skills.
- Ability to collaborate with geographically distributed teams and perform effectively in high-pressure production environments.
Benefits
- Competitive compensation including base pay and annual incentive.
- Comprehensive health and life insurance and well-being benefits based on location.
- Pension and retirement benefits.
- Paid time off, personal/family care, and other leaves of absence.
- Flexible hybrid work model with three days onsite and two days remote, including onsite Tuesdays, Wednesdays, and a third team- or employee-specific day.
Tech Stack
Apache KafkaAWSBashCOBOLGrafanaIBM DB2JavaScriptLinuxOpenShiftPerlPostgreSQLPythonRabbitMQRubySeleniumSnowflakeSplunkSQLWindows
Categories
Site Reliability