SRE

ALPHARETTA, GA 30005

Posted: 03/30/2021 Employment Type: Contract To Hire Category: Site Reliability Engineer Job Number: 55218

Job Description


Looking for a highly motivated Site Reliability Engineer, who is capable of build and run large-scale, massively distributed, fault-tolerant systems. Individual to work with teams across the organization and ensures core services reliability and keep an eye on capacity and performance.

 

Responsibilities:

• Responsible for blameless postmortems and proactive identification of potential outages factor into iterative improvement.

• Experience in Designing and Deploying multi-data center Large Scale Web Applications.

• Work closely with dev, and ops teams to build highly available, cost effective systems.

• Create new tools and scripts designed for auto-remediation of incidents.

• Design/Implementation of Big Data technologies, including Hadoop, MongoDB, Kafka, RabbitMQ, Zookeeper, Spark, ELK, etc

• Responsible for establishing end-to-end monitoring and alerting on all critical aspects to ensure SLAs and get proactive notifications of possible issues for all systems.

• Design platforms for extremely high uptime metrics.

• Works well independently and requires little or no supervision.

• Work with cloud operations team to resolve trouble tickets, developing and running scripts, and troubleshooting.

• Fully understand the application, microservices interactions.

• Design/Implementation containers/applications in scalable HA/DR multi-tier cloud environments, including new system design, documentation, implementation, and deployment.

Job Requirements:

• 7+ years of experience in the following areas:

• Experience in providing L4 technical support for production 24x7.

• Strong experience in production support and operations.

• Design/Implementation of network and presentation tier technologies, including F5, Apache, Nginx, etc

• Experience in Performance Testing/Tuning/Monitoring, maximizing system uptime and availability, ensuring functional and performance SLAs.

• Experience with monitoring Application/Infrastructure Performance, and availability.

• Automation Experience with Build/deployment, Software Configuration/Continuous Integration/Continuous Delivery/Release Engineering related tasks in an JavaEE/C++ Environments.

• Experience in automating manual processes using Python, Ruby, Unix Shell (bash, ksh), perl, Ant, etc.

• Installing, Configuring, Administering, and Tuning of JavaEE Application Servers/Containers like Tomcat, WebSphere, etc

• Installing/maintaining/Administering software on Unix Linux, Windows servers.

• Experience with Web service technologies, including REST, SOAP, JSON, XML

• Experience with Cloud Platforms and virtualization Technologies.

• Deploying and automating infrastructure/applications in cloud environment using Chef, RPM, etc.

• Working closely with Development, QA, Product Management, and Production Ops teams to make sure Product Releases on-time with quality.

• Hands on experience Configuring and Administering SCM(GIT, SVN), Build (CMake, Make files, Maven), CI(Jenkins), CD Automation Tools.

• Experience with database (RDBMS, NoSql) technologies is a plus.
Apply Online
Apply with LinkedIn Apply with Facebook Apply with Twitter

Send an email reminder to:

Share This Job:

Related Jobs:

Login to save this search and get notified of similar positions.