JOB DESCRIPTION
We are looking for a DevOps / System Operation Engineer to join our team and take responsibility for operating, monitoring, troubleshooting, and maintaining the Britymail Service.
The successful candidate will work closely with Infrastructure, Development, and other technical teams to ensure system stability, availability, performance, and smooth deployment. The role requires hands-on experience with AWS/On-Premise environments, Kubernetes, Docker, CI/CD, Web/WAS servers, system monitoring, and incident management.
A. Britymail Service Operation
Operate and maintain the Britymail Service in accordance with established operation procedures.
Understand and follow Jira-based operation processes, including:
System overview and operation processes.
Release management and deployment procedures.
Regular operation processes and documentation.
Perform daily system checks and monitor outstanding/unclosed system events.
Report and follow up on system issues until they are properly resolved.
Prepare detailed analysis reports covering:
Incident symptoms and impact.
Corrective and preventive actions.
Investigate system incidents, identify root causes, and propose appropriate solutions.
Review team members' tasks and provide technical support when required.
B. System Monitoring & Incident Management
Monitor system health, availability, and performance.
Monitor infrastructure resources, including:
Other system resources as required.
Detect, analyze, and troubleshoot system incidents.
Collaborate with Infrastructure and other technical teams to resolve system issues.
Request, coordinate, and verify resource installation and configuration.
Review resource installation plans and provide technical feedback.
Follow established incident management and escalation processes.
C. Kubernetes & Container Operations
Deploy, manage, and monitor containers in Kubernetes (K8S) environments.
Manage Kubernetes workloads and resources.
Perform application deployment and troubleshooting on Kubernetes.
Work with Helm for application packaging and deployment.
Manage and troubleshoot Docker containers.
Monitor container health, resource usage, and operational status.
Investigate and resolve deployment or runtime issues in containerized environments.
D. Web / WAS Server Operation
Install, configure, deploy, and maintain Web servers such as:
Install, configure, deploy, and maintain WAS servers such as:
Analyze and troubleshoot Web/WAS-related issues.
Investigate application deployment, connectivity, performance, and runtime issues.
E. DevOps & Middleware Operation
Support and maintain CI/CD pipelines.
Troubleshoot issues related to automated build, deployment, and release processes.
Develop and maintain Shell scripts for system operation and automation.
Operate and troubleshoot Message Queue systems such as:
Support NoSQL platforms such as:
Collaborate with Development and Infrastructure teams to improve system reliability and operational efficiency.
REQUIREMENTS
Hands-on experience with AWS Cloud and/or On-Premise environments.
Experience in system operation and monitoring.
Strong practical experience with:
Web Server (Nginx / Apache)
Experience in system incident investigation and troubleshooting.
Understanding of incident management and escalation processes.
Experience working with system monitoring and operation tools.
Ability to coordinate and work effectively with Infrastructure and Development teams.
Good problem-solving and troubleshooting skills.
Experience with NoSQL databases, especially:
Experience with monitoring and operation tools such as:
Basic software development/programming skills.
Experience with system automation and operational optimization.
Good understanding of DevOps practices and methodologies.
Good written and verbal communication skills in English.
BENEFITS
Minimum 13 months salary per year - not including other bonuses such as KPI bonus for work efficiency, project bonus and revenue bonus. We do performance review twice a year. You will work in a professional, dynamic and friendly environment.
VTI offer annual health check-ups and fully pay social insurance, health insurance and unemployment insurance premium following the Labor law.
We offer one vacation/company trip and 4 teambuilding trips per year for every employees, along with various entertainment activities including: Swimming Clubs, Yoga, Zumba, Kendo and music order via internal Radio channel.
Two MVPs will be rewarded with a free trip to Japan, Taiwan, Singapore, or else.
We offer variety promotion opportunities and chance to raising income for people with capacity, enthusiasm, and long-term commitment.
We offer free Japanese class at the company.
We provide training opportunities to help our people to improve their skills. We support our members to learn and get Cloud, AWS, PMF, PMP certification.
Working hour from: 08:30am to 05:30pm. From Monday to Friday.