Ops Lead Engineer (#554)
¥7,500,000 ~ ¥9,500,000 Yearly
Apply港区白金, 東京都
Full time Permanent
Insurance
Job description
日本語版 (Japanese Version)
Ops Engineerは、ビジネスアプリケーションの運用・リリース・自動化に携わるプロフェッショナルエンジニアです。定められた標準やプロセスに基づき、日々の運用を安定して遂行しながら、CI/CD、自動化、監視の改善に前向きに取り組んでいただきます。Tribe / Squad の主担当Opsエンジニアとして、プロジェクトと日常運用(BAU)の双方に関わります。
職務内容 / Responsibilities
-
ビジネスアプリケーションの運用・保守: インシデント対応(一次・二次対応を中心とした調査・復旧)、エスカレーション時の状況整理。既存のRunbook / KEDBを活用した切り分け、および必要に応じた改善提案。
-
変更・リリース作業の実行: Jenkins、Control-M、GitHubなどを利用したテスト/本番環境へのデプロイ作業。変更・リリースプロセス(チケット起票、承認フロー、ロールバック手順)の遵守、およびOpsの観点からのリリース準備作業(事前チェック、手順確認など)への参画。
-
CI/CD・自動化の実装サポート: Jenkins、Jira、GitHubなどを用いたパイプラインの実装・保守。日常的な運用作業(デプロイ、データ抽出、バッチ実行など)の自動化スクリプト作成・改善、および上位エンジニアの方針に基づく小〜中規模な改善の実装。
-
監視・可観測性の整備: Prometheus、Grafana、Dynatrace、CloudWatch、Splunkなどを用いたダッシュボード/アラート設定。アラートの見直し・調整を通じた検知精度向上およびMTTR短縮への貢献。
-
インフラ・プラットフォーム運用: AWS、OpenShift、Managed Public IaaS上でのアプリケーション運用・設定変更の実施。ネットワーク(DNS, LB, FW)や証明書更新など、基盤チームと連携した基本的な運用作業。
-
ナレッジ共有とチーム貢献: Runbook / KEDB / Wiki の作成・更新(対応したインシデントや作業内容の整理)。組織で採用するアジャイル開発への参画、スプリントやカンバンを用いたタスク管理とチーム内共有。チームメンバーやパートナーへの情報共有・技術的サポート。
-
ステークホルダーとの協働: プロダクトチームや開発チームと連携し、リリース計画・品質改善・パフォーマンス改善に貢献。必要に応じた日本語/英語でのコミュニケーション(メール・チャット・ミーティング)。
求めるスキル・経験 / Required Skills and Experiences
【Must(必須要求)】
-
経験・マインドセット:
-
IT運用/SRE/DevOps いずれかの領域での実務経験 おおよそ2年以上
-
運用・トラブルシューティングに前向きに取り組み、原因を理解しようとする姿勢
-
与えられた要件・タスクに対し、必要な技術的な手段を自ら考えて実行できること
-
チームでのコラボレーションや情報共有を大切にできる方
-
-
技術スキル:
-
Linux(RHEL系)または Windows Server の運用経験
-
Jenkins などのCIツールを使ったビルド/デプロイの実務経験
-
Git を用いたソースコード管理(ブランチ運用・Pull Requestなど)の実務経験
-
いずれかのスクリプト/言語での自動化経験(例:Shell, Python, Groovy, PowerShell, Node.js など)
-
AWSやコンテナ(Docker / Kubernetes / OpenShift)の基本的な理解・利用経験
-
Prometheus, Grafana, CloudWatch, Splunk, Dynatrace 等のいずれかを用いた監視/ログ分析の経験
-
-
プロセス・働き方:
-
CI/CDやGitフローなど、基本的なソフトウェアエンジニアリングプラクティスへの理解
-
Agile(Scrum / Kanban)でのチーム開発、またはJira等によるタスク管理の経験
-
手順書/Runbook/Wiki 等のドキュメントを、指示に基づき自ら作成・更新できること
-
-
コミュニケーション:
-
日本語:ビジネスレベルが望ましい(社内会話・基本的なドキュメントが可能であること。※チーム構成や候補者の強みに応じて柔軟に検討可能)
-
英語:技術ドキュメントの読解に抵抗がないこと(メール・チャットでの簡単なやり取りができれば尚可)
-
【Nice to have(歓迎要求)】
-
金融・保険業界のシステム運用経験
-
Control-M等のジョブスケジューラの運用経験
-
IaC(Terraform等)や構成管理ツールの利用経験
-
SREプラクティス(SLO / エラーバジェット / ポストモーテム)への興味・経験
-
他クラウド(Azure, GCP)の知識・経験
-
海外チーム/ベンダーとの協働経験
English Version
The Ops Engineer is a professional engineer responsible for the operations, release management, and automation of business applications. Under established standards and processes, you will help ensure stable day-to-day operations while proactively working on CI/CD, automation, and monitoring improvements. You will act as the primary Ops engineer for one or more Tribes/Squads, contributing to both project deliveries and BAU (Business As Usual) activities.
Key Responsibilities
-
Operations and Maintenance: Handle incident support (mainly L2 investigation and recovery), including clear documentation and situational synthesis during escalations. Perform initial analysis using existing runbooks / KEDB and suggest operational improvements where appropriate.
-
Change and Release Execution: Execute deployments to test and production environments using Jenkins, Control-M, GitHub, and related tools. Strictly follow defined change and release processes (ticket creation, approval flows, rollback procedures) and support release readiness activities from an Ops perspective (pre-checks, procedure validation, etc.).
-
CI/CD and Automation Support: Implement and maintain CI/CD pipelines using Jenkins, Jira, GitHub, and related tools under the guidance of senior engineers. Develop and improve automation scripts for routine operational tasks (deployments, data extraction, batch execution, etc.) and implement small to medium-sized technical enhancements.
-
Monitoring and Observability: Configure comprehensive dashboards and alerting frameworks using Prometheus, Grafana, Dynatrace, CloudWatch, Splunk, and other observability tools. Review and tune alerts to improve signal quality, eliminate noise, and contribute to MTTR reduction.
-
Infrastructure and Platform Operations: Support application operations and configuration changes on AWS, OpenShift, and Managed Public IaaS. Perform basic operational infrastructure tasks in collaboration with core platform teams, including DNS, load balancer, firewall settings, and SSL/TLS certificate updates.
-
Knowledge Sharing & Agile Contribution: Create and update runbooks, KEDB, and wiki pages based on resolved incidents and completed operational tasks. Participate in the organization’s Agile development practices, using sprints or Kanban for task management, and provide technical support to team members and partners where needed.
-
Stakeholder Collaboration: Work closely with Product and Engineering teams to support release plans, quality improvements, and performance tuning. Communicate fluently in Japanese and/or English as needed across emails, chats, and cross-functional meetings.
Required Skills and Experiences
【Must】
-
Experience & Mindset:
-
Approximately 2+ years of hands-on experience in IT Operations, SRE, DevOps, or related technical domains.
-
A positive, analytical attitude toward operations and troubleshooting, with a passion for root-cause analysis.
-
Ability to interpret assigned requirements and independently determine appropriate technical approaches for execution.
-
Strong commitment to teamwork, transparency, and proactive knowledge sharing.
-
-
Technical Skills:
-
Experience operating Linux (RHEL family) and/or Windows Server environments.
-
Practical experience using CI tools such as Jenkins for building and deploying applications.
-
Solid proficiency with source code management using Git (branching strategies, pull requests, etc.).
-
Automation experience in at least one scripting or programming language (e.g., Shell, Python, Groovy, PowerShell, Node.js).
-
Basic understanding and practical usage experience of AWS and container orchestration technologies (Docker, Kubernetes, OpenShift).
-
Hands-on experience with monitoring and log analysis using at least one major tool: Prometheus, Grafana, CloudWatch, Splunk, Dynatrace, etc.
-
-
Process & Ways of Working:
-
Strong understanding of modern software engineering practices, including CI/CD pipelines and Git flows.
-
Experience working within Agile frameworks (Scrum / Kanban) and utilizing Jira or similar tools for sprint task management.
-
Ability to independently draft, update, and maintain comprehensive technical documentation (procedures, runbooks, wikis).
-
-
Communication:
-
Japanese: Business-level preferred (capable of driving internal meetings and compiling basic documentation). Language requirements can be flexibly discussed based on the overall team composition and technical strength.
-
English: Professional reading comprehension of technical documentation; basic written communication (emails/chats) is highly desired.
-
【Nice to Have】
-
Systems operations experience within the financial or insurance industries.
-
Experience utilizing enterprise job schedulers such as Control-M.
-
Familiarity with Infrastructure as Code (IaC) tools (e.g., Terraform) or configuration management systems.
-
Active interest or practical experience in SRE practices (SLOs, error budgets, post-mortems).
-
Knowledge or operational experience with other public cloud providers (Azure, GCP).
-
Prior experience collaborating closely with overseas technical teams or offshore vendors.
Language requirement
Working hours
Back to jobs