EOS RPO

Staff, Software Engineer

Posted May 21, 2026
Location
Chennai, Tamil Nadu
Hours/week
45 hrs/week

Requirements:

Architecture Acumen: Requires knowledge of: Architectural principles; Systems and environment behavior; Architectural Styles, Patterns and plans; Architectural standards; Non-functional System performance parameters; Technology Strategy.

Defect Management and Troubleshooting: Requires knowledge of: Defect life-cycle process, defect tracking tools and methodologies; Defect reporting; Regression testing; Root cause analysis; Root cause corrective action. To conduct root cause analysis (RCA) and root cause corrective action (RCCA) to identify the origin of defects/ performance gaps and prevent them from recurring. Track registered issues for the product/solution and prioritize them for resolution. Measure usability of the product/solution as per customer/business requirement after defect fixing and plugging test gaps. Analyze the issues and plan a series of steps which potentially includes reconfiguration, integration, removal or addition of application components to enhance the application's functionality, usability and security. 

DevOps Orientation: Requires knowledge of: Different operating systems; Software maintenance tools and techniques; Application monitoring tools and techniques; Debugging tools; Mock screen; Pseudocodes; Reverse Engineering; Traceability matrix; System performance, security, integration; Data migration and accessibility; Design Methodologies.

Agentic AI framework—requires a shift from passive monitoring to active, autonomous oversight of digitalisation

 

The core requirement is managing "human-in-the-loop" systems where AI agents plan and execute tasks, while the TDO ensures alignment with business goals, security, and safety.

Prompt Engineering and Evaluation: Proficiency in designing prompts that enable agents to act reliably and setting up evaluation frameworks for monitoring performance.

Requirement And Scoping Analysis: Requires knowledge of: Traceability matrix; Risk analysis methodologies; Cost Analysis; Business objectives; Classification of requirements; User stories To explore relevant products/solutions from an existing repertoire, that can address business/technical needs.

  • Strong and demonstrable incident management skills with relevant experience in an enterprise organization.

  • Methodical and systematic problem solving approach, combined with a solid awareness of ownership, initiative and drive.

  • Experience investigating, analysing and troubleshooting large scale enterprise systems.

  • Understanding of Unix/Linux systems from kernel to shell and beyond, taking in system libraries, file systems, and client-server protocols along the way.

  • Experience working with and developing enterprise monitoring/tooling solutions like Grafana, Prometheus, Kibana, Splunk, Graphite, Dynatrace, catchpoint.

  • Working knowledge of one or more cloud technologies such as AZURE, GCP and OpenStack.

  • Expert verbal and written communication skills.

  • Demonstrate excellent judgement in decision making.

  • Strong focus on collecting and inferring metrics.

  • Excellent communication skills


What you’ll bring:

  • 10-14 years in an infrastructure, systems, engineering or development environment delivering operational excellence to highly complex distributed systems.

  • Bachelor's Degree in Computer Science or a related field, or relevant work experience of 10+ years.

  • Experience and exposure working in a 24/7 operations support environment.

  • Working and technical expertise in K8 and microservice architectures.

  • Experience administering Unix/Linux in a production environment.

  • Ability to supervise the Site Reliability Operations team, mentor and provide guidance.

  • Working knowledge of BASH, Python, AI or other scripting languages

  • Utilize AI-powered monitoring and anomaly detection tools to predict potential failures and resource bottlenecks before they impact users.

  • Ensure the reliability, performance, and scalability of infrastructure specifically designed for AI/ML workloads,

  • Networking knowledge and understanding of network concepts, such as different protocols (TCP/IP, UDP, ICMP, etc.), MAC addresses, IP packets, DNS, OSI layers, and load balancing).

Similar jobs

+ Search all jobs