Operations Lead Engineer (#537)


Tokyo
Full time Permanent
Insurance

Job description

We are seeking an Operations Lead Engineer to drive the reliability, release automation, and continuous improvement of our business applications. This role balances day-to-day production stability with proactive SRE/DevOps initiatives in a highly collaborative, global environment.

Key Responsibilities

1. Production Operations & Release Management

  • L1/L2 Support: Lead incident response, troubleshooting, and root cause analysis (KEDB).

  • Release & Batch Execution: Coordinate change management and deployment activities using Jenkins, GitHub, and Control-M.

2. CI/CD & Automation (SRE)

  • Pipeline Engineering: Design, implement, and optimize CI/CD pipelines (Jenkins, GitHub Actions).

  • Toil Reduction: Automate repetitive tasks (deployments, data extraction, batch jobs) using scripting languages.

3. Monitoring & Observability

  • Alerting & Metrics: Build and maintain telemetry dashboards using Prometheus, Grafana, Dynatrace, CloudWatch, and Splunk.

  • MTTR Reduction: Optimize alerting logic to accelerate incident recovery and analyze performance trends.

4. Cloud & Platform Collaboration

  • Hybrid Cloud Support: Manage application configurations on AWS and OpenShift.

  • Infrastructure Coordination: Collaborate on network configs (DNS, Load Balancers, Firewalls) and SSL certificate renewals.

Required Skills & Experience

Must Have

  • Experience: 3+ years in IT Operations, SRE, or DevOps.

  • OS & Cloud: Strong administration skills in Linux (RHEL) or Windows Server; basic knowledge of AWS and containerization (Docker, Kubernetes/OpenShift).

  • CI/CD & Git: Solid experience with Jenkins and Git-flow (branch management, PRs).

  • Scripting: Proficiency in at least one scripting language (Python, Shell/Bash, Groovy, PowerShell).

  • Observability: Hands-on experience with modern monitoring tools (e.g., Grafana, Prometheus, Splunk).

  • Agile & Docs: Experience working in Scrum/Kanban (Jira) and writing technical runbooks.

Nice to Have

  • Experience in the Financial or Insurance industries.

  • Experience with enterprise job schedulers (like Control-M).

  • Knowledge of Infrastructure as Code (IaC) via Terraform.

  • Practical understanding of SRE concepts (SLOs, Error Budgets, Post-mortems).

Language requirement

Japanese (Fluent), English (Fluent)

Working hours

9:00-18:00

Back to jobs