NHS Human Services, Inc.

Mobile nhs-human-services Logo

Job Information

Cribl, Inc Senior Site Reliability Engineer in Jackson, Mississippi

This is a Job Description for a Senior Site Reliability Engineer in Jackson, Mississippi

Summary: We are looking for Cloud Site Reliability Engineers and Developers at all levels at Cribl, who enjoy being in the thick of it. Fixing things at the operational side should always be the last resort, so our SRE engineers are involved from conception to design to development and all the way through production and beyond. You provide your creative input into all things Cloud, Scaling, Reliability, High Availability and much more.

Duties & Responsibilities:

Engage with teams and improve service delivery and reliability across their entire lifecycle. Measure and monitor all production systems with an eye towards availability, latency, and overall system health. Seek out the cause of errors and instability in our production cloud services and drive teams towards better operational excellence. Engage with product and platform teams to improve and evolve systems by lobbying for changes that improve reliability, resilience, and observability. Help Identify and drive down toil with creative innovation and automation. On-call responsibilities

 

Requirements

and Qualifications[]{#Hlk142289191}[]{#Hlk142304825}

:

 Extensive experience with enterprise scale continuous delivery environments. 5+ years of experience as a DevOps or SRE. Development with JavaScript/Node.js/TypeScript in a Linux/Mac environment. Experience with Configuration Management Tools like Terraform (preferred) or Puppet, Chef, Ansible. Experience with sustainable incident response in a blameless environment. Knowledge of cloud platforms (prefer Azure) and container + orchestration technologies. Experience with APM and Observability and related tools such as, New Relic, Splunk, CloudWatch, Prometheus, Grafana/Kibana, Sentry etc. Background in Linux Systems Engineering. Experience with Incident response related tools for instance, PagerDuty, Fire Hydrant, Blameless etc. Comfortable with a high level of autonomy and working with a distributed team

Equal Opportunity/Affirmative Action Employer.

 

DirectEmployers