
Site Reliability Engineer Intern
Copart, Inc.2 hours ago
Dallas, TX, USAIntern
Responsibilities
- Monitor global data centers and application infrastructure and proactively resolve issues before they affect operations.
- Design, build, and optimize monitoring and automation tools using Python and Ansible for metrics collection and automated remediation.
- Maintain monitoring, collaboration, and Kubernetes-related observability tooling, including Datadog.
- Perform monthly security patching for operating systems and critical applications.
- Troubleshoot and communicate infrastructure and application issues with cross-functional teams.
- Develop performance reporting and quality control capabilities using Datadog.
- Create and maintain standard operating procedures, system diagrams, and training materials.
- Support incident triage, tracking, impact assessment, severity classification, root cause analysis, and post-incident reviews.
- Coordinate and perform periodic failover testing for network, systems infrastructure, and application environments.
- Identify recurring issues and opportunities for preventive action and process improvement.
- Process operational change paperwork and support 24/7 operations.
Requirements
- Experience triaging, tracking, and resolving incidents in a production environment.
- Ability to perform initial impact assessments and severity classifications and communicate updates during active incidents.
- Experience documenting incident timelines, actions, and resolution details.
- Core knowledge of Linux and Windows, with familiarity with virtual environments, basic networking, scripting, automation, and observability tools.
- Intermediate proficiency in programming and scripting is highly desired.
- Strong troubleshooting skills across Linux, Unix, and Windows-based systems are highly desired.
- Familiarity with Datadog is a plus.
- Hands-on experience with virtual machine management software, particularly VMware vSphere, is a plus.
- Basic understanding of AI tools and concepts, including AI-assisted workflows, prompt-based tools, or automation-enhanced support processes, is a plus.
- Experience with AWS or GCP is a plus.
- Strong collaborative, interpersonal, oral, and written communication skills and the ability to work independently and manage competing priorities.
Benefits
- The role is based at Copart’s Dallas location and supports a flexible work schedule as part of 24/7 operations.
- Copart emphasizes diversity, inclusion, collaboration, and opportunities to grow and contribute meaningfully.
Tech Stack
Categories
Site Reliability