Required Skills: What you will need to be successful in this role:
• Strong Linux expertise is a must.
• Expert-level skills and background in systems administration and engineering.
• Understanding of the components of a cloud infrastructure, including hardware platforms, OS, applications, databases, networks, web and application servers.
• Prior experience in Site Reliability Engineering/DevOps and managing large-scale server infrastructure in a cloud computing or MSP setting is highly desirable.
• Experience with performance and availability monitoring, analysis, and configuration management platforms (e.g., Nagios/Icinga, Cacti, Ansible, Puppet, cfengine, chef, Splunk, Logstash).
• Working-level knowledge of one: Perl, Python, JavaScript.
• Familiarity with MySQL, Oracle, MariaDB, PostgreSQL or similar technologies; proficiency preferred.
• Expert-level skills and experience with service troubleshooting in a production environment covering web front-end, Systems, databases, and Networks.
• Familiarity with Networking Technologies such as routing, switching, and load balancing. F5 and NGINX experience is ideal.
• Understanding of ITIL v3 framework and how it applies to incidents, problems and changes.
• Good communication skills and ability to work well in a collaborative team environment
Job Overview:
Duties: The Team:
As a key member of the Systems Administration team within Operations Engineering, you will be responsible for the administration and operations of the global cloud infrastructure that runs our SaaS product. This is an opportunity to be at the core of running a Cloud SaaS platform that scales to millions of users! The Cloud Operations team is responsible for ensuring the availability and efficiency of the server infrastructure that runs our SaaS platform while consuming and deploying products that have been newly developed by engineering teams. You will be working closely with engineers and developers across the company.
What you will do in this role:
• Contribute to Configuration Management and Infrastructure as Code for the client's global private cloud.
• Develop tools in Python, bash, and JavaScript to replace manual work and improve customer maintenance experience.
• Drive enhancements and bug fixes for large-scale automation projects such as patching, provisioning, and kickstart domains.
• Design and implement procedures to accomplish maintenance where automation and tooling cannot; drive resolution of root causes with internal team members.
• Prepare new Client products and services for production readiness with design review, feedback to engineering teams, training, and testing.
• Use broad knowledge and experience of systems administration and networking principles to proactively prevent and address incidents while constantly improving documentation.
• Participate in escalations and Root Cause Analysis of issues in both US Federal and global Commercial infrastructures.
• Troubleshoot database backup and restore failures as well as perform database migrations.
• Support operation of a wide variety of infrastructure services including Machine Learning and Prediction, Cloudera Big Data clusters, Kafka and RabbitMQ messaging, database encryption, E-Mail infrastructure at scale, DNS, Puppet, Elasticsearch, F5 BigIP, and more.
- **Only those lawfully authorized to work in the designated country associated with the position will be considered.**
- **Please note that all Position start dates and duration are estimates and may be reduced or lengthened based upon a client’s business needs and requirements.**
I am very happy with the Rose International, and the professionalism of the employees.
Robin, Consultant
As a contractor, I have to say that Rose International was by far the best agency I have worked for.
Q'testdalir, Consultant
You are customer service oriented. No matter whether it was the Recruiter or someone in Human Resources/Payroll, you were responsive. That to me is key!
Tonya, Consultant
Each time I contacted Rose, I was completely satisfied with the great attention and customer service I received. Each person was extremely knowledgeable and patient with my concerns or questions.
Diana, Consultant
Any time I did have a question and called, the phone was always answered, and my question/concern was immediately resolved.
Sally, Consultant
EMPLOYEE COMMENTS