EOS RPO
Staff, Software Engineer
Requirements:
Architecture Acumen: Requires knowledge of: Architectural principles; Systems and environment behavior; Architectural Styles, Patterns and plans; Architectural standards; Non-functional System performance parameters; Technology Strategy.
Defect Management and Troubleshooting: Requires knowledge of: Defect life-cycle process, defect tracking tools and methodologies; Defect reporting; Regression testing; Root cause analysis; Root cause corrective action. To conduct root cause analysis (RCA) and root cause corrective action (RCCA) to identify the origin of defects/ performance gaps and prevent them from recurring. Track registered issues for the product/solution and prioritize them for resolution. Measure usability of the product/solution as per customer/business requirement after defect fixing and plugging test gaps. Analyze the issues and plan a series of steps which potentially includes reconfiguration, integration, removal or addition of application components to enhance the application's functionality, usability and security.
DevOps Orientation: Requires knowledge of: Different operating systems; Software maintenance tools and techniques; Application monitoring tools and techniques; Debugging tools; Mock screen; Pseudocodes; Reverse Engineering; Traceability matrix; System performance, security, integration; Data migration and accessibility; Design Methodologies.
Agentic AI framework—requires a shift from passive monitoring to active, autonomous oversight of digitalisation.
The core requirement is managing "human-in-the-loop" systems where AI agents plan and execute tasks, while the TDO ensures alignment with business goals, security, and safety.
Prompt Engineering and Evaluation: Proficiency in designing prompts that enable agents to act reliably and setting up evaluation frameworks for monitoring performance.
Requirement And Scoping Analysis: Requires knowledge of: Traceability matrix; Risk analysis methodologies; Cost Analysis; Business objectives; Classification of requirements; User stories To explore relevant products/solutions from an existing repertoire, that can address business/technical needs.
Strong and demonstrable incident management skills with relevant experience in an enterprise organization.
Methodical and systematic problem solving approach, combined with a solid awareness of ownership, initiative and drive.
Experience investigating, analysing and troubleshooting large scale enterprise systems.
Understanding of Unix/Linux systems from kernel to shell and beyond, taking in system libraries, file systems, and client-server protocols along the way.
Experience working with and developing enterprise monitoring/tooling solutions like Grafana, Prometheus, Kibana, Splunk, Graphite, Dynatrace, catchpoint.
Working knowledge of one or more cloud technologies such as AZURE, GCP and OpenStack.
Expert verbal and written communication skills.
Demonstrate excellent judgement in decision making.
Strong focus on collecting and inferring metrics.
Excellent communication skills
What you’ll bring:
10-14 years in an infrastructure, systems, engineering or development environment delivering operational excellence to highly complex distributed systems.
Bachelor's Degree in Computer Science or a related field, or relevant work experience of 10+ years.
Experience and exposure working in a 24/7 operations support environment.
Working and technical expertise in K8 and microservice architectures.
Experience administering Unix/Linux in a production environment.
Ability to supervise the Site Reliability Operations team, mentor and provide guidance.
Working knowledge of BASH, Python, AI or other scripting languages
Utilize AI-powered monitoring and anomaly detection tools to predict potential failures and resource bottlenecks before they impact users.
Ensure the reliability, performance, and scalability of infrastructure specifically designed for AI/ML workloads,
Networking knowledge and understanding of network concepts, such as different protocols (TCP/IP, UDP, ICMP, etc.), MAC addresses, IP packets, DNS, OSI layers, and load balancing).