Operations Lead Engineer (#537)
Tokyo
正社員
保険
求人概要
We are seeking an Operations Lead Engineer to drive the reliability, release automation, and continuous improvement of our business applications. This role balances day-to-day production stability with proactive SRE/DevOps initiatives in a highly collaborative, global environment.
Key Responsibilities
1. Production Operations & Release Management
-
L1/L2 Support: Lead incident response, troubleshooting, and root cause analysis (KEDB).
-
Release & Batch Execution: Coordinate change management and deployment activities using Jenkins, GitHub, and Control-M.
2. CI/CD & Automation (SRE)
-
Pipeline Engineering: Design, implement, and optimize CI/CD pipelines (Jenkins, GitHub Actions).
-
Toil Reduction: Automate repetitive tasks (deployments, data extraction, batch jobs) using scripting languages.
3. Monitoring & Observability
-
Alerting & Metrics: Build and maintain telemetry dashboards using Prometheus, Grafana, Dynatrace, CloudWatch, and Splunk.
-
MTTR Reduction: Optimize alerting logic to accelerate incident recovery and analyze performance trends.
4. Cloud & Platform Collaboration
-
Hybrid Cloud Support: Manage application configurations on AWS and OpenShift.
-
Infrastructure Coordination: Collaborate on network configs (DNS, Load Balancers, Firewalls) and SSL certificate renewals.
Required Skills & Experience
Must Have
-
Experience: 3+ years in IT Operations, SRE, or DevOps.
-
OS & Cloud: Strong administration skills in Linux (RHEL) or Windows Server; basic knowledge of AWS and containerization (Docker, Kubernetes/OpenShift).
-
CI/CD & Git: Solid experience with Jenkins and Git-flow (branch management, PRs).
-
Scripting: Proficiency in at least one scripting language (Python, Shell/Bash, Groovy, PowerShell).
-
Observability: Hands-on experience with modern monitoring tools (e.g., Grafana, Prometheus, Splunk).
-
Agile & Docs: Experience working in Scrum/Kanban (Jira) and writing technical runbooks.
Nice to Have
-
Experience in the Financial or Insurance industries.
-
Experience with enterprise job schedulers (like Control-M).
-
Knowledge of Infrastructure as Code (IaC) via Terraform.
-
Practical understanding of SRE concepts (SLOs, Error Budgets, Post-mortems).
必須言語
勤務時間
求人票に戻る