About Interpay
Interpay builds the retail payment software that banks and merchants across Saudi Arabia depend on: POS and SoftPOS acceptance, e-commerce gateways, payment switching and merchant services. Our products handle cardholder data under SAMA regulation and PCI security standards, and they run around the clock.
The role
Interpay's payment platform runs 24/7 across Linux and Windows servers, Kubernetes clusters, SQL Server and YugabyteDB databases, and the network links that connect terminals, merchants and banks. When an issue is beyond Level 1 and Level 2, it lands with Level 3. You'll find the root cause across servers, network, databases, application configuration and Kubernetes, fix what can be fixed within support, and prevent recurrence. When an issue needs a code change, you'll coordinate the fix with development and remain the owner until it's closed. You'll also execute Kubernetes deployments and, when necessary, roll them back. You'll report to the DevOps Manager.
What you'll do
Escalation and incident resolution
- Level 3 ownership: own escalations from Level 2 that need server, network, database, configuration or platform expertise, through to permanent resolution
- Root cause analysis: analyse system, application, database and Kubernetes logs, and confirm the cause with evidence
- Fault repair: fix directly within your authority, or define the fix precisely and coordinate it with DevOps
- Development coordination: escalate code issues with a reproducible case, full diagnostics and an impact assessment, and own them until they're closed with the customer
- Problem management: document root cause analyses and drive permanent fixes through change control
- Major incidents: lead the technical response, restore service and write post-incident reports
Servers and network
- Support Linux and Windows Server: services, processes, storage, file systems, scheduled jobs, users and permissions
- Tune CPU, memory, disk and I/O, and resolve capacity issues before they affect service
- Support OS patching, hardening and vulnerability remediation in line with PCI DSS
- Manage TLS certificates and service accounts so renewals and expiries never cause outages
- Troubleshoot TCP/IP, DNS, routing, NAT, firewalls, load balancers and VPN or leased links to banks, schemes and processing hosts
- Diagnose terminal-to-host and host-to-host problems, from connection timeouts to TLS handshake failures, down to packet level with tcpdump and Wireshark
- Specify and verify firewall rules, ports and access paths with the network and security teams
Databases and configuration
- Monitor SQL Server and YugabyteDB for availability, replication, performance, storage growth, locking and long-running queries
- Support YugabyteDB in an active-active, multi-datacentre setup, including node health, replication status and failover
- Write SQL to check data integrity and catch orphaned records, stuck or duplicated transactions and reconciliation mismatches
- Apply approved data corrections under change control, with backups and audit trail, and verify backups and restores
- Know Interpay's configuration files and tables, find misconfiguration and environment drift, and keep production, DR and test consistent
Kubernetes operations and deployment
- Support clusters, nodes, pods, deployments, services, ingress, persistent storage, ConfigMaps and Secrets
- Fix crash loops, failed scheduling, resource exhaustion, image pull errors, networking, DNS and storage issues
- Execute deployments strictly to Interpay's procedure, with pre-deployment checks and post-deployment verification
- Roll back promptly and safely when a release degrades service, restore the last known-good state, and document the event
- Use metrics, logs and dashboards to catch degradation before it becomes an outage
Knowledge and compliance
- Write runbooks and known-error records so Level 1 and Level 2 resolve more without escalation
- Coach Level 2 engineers in log analysis, database investigation and platform diagnostics
- Follow PCI DSS, least privilege and segregation of duties, and keep card numbers, PINs and keys out of logs, tickets and query outputs
- Support internal audits, penetration tests and certification, and remediate assigned findings
What you'll bring
- Bachelor's degree in Computer Science, IT, Engineering or a related field, or equivalent practical experience
- 5+ years in technical support, systems administration, application support or site reliability, including 2+ years supporting production systems at Level 3 or an equivalent tier
- Linux and Windows Server: services, storage, permissions, scheduling, patching and performance troubleshooting
- Networking: TCP/IP, DNS, routing, firewalls, load balancing and TLS, with hands-on connectivity diagnosis
- Monitoring SQL Server and a distributed SQL database such as YugabyteDB or PostgreSQL, including replication and performance health
- Strong SQL, and a proven ability to check data integrity and detect anomalies
- Root-cause log analysis on production systems
- Application configuration files and tables, and configuration-related faults
- Kubernetes: finding and fixing faults in running clusters, and executing deployments and rollbacks to a defined procedure
- Formal change management in a production environment
- Calm, methodical work under incident pressure, and clear root cause and post-incident reports
- Professional working proficiency in English
- Willing to join an on-call rotation and work outside standard hours for incidents and planned maintenance
- Willing and eligible to work on-site in Riyadh
Nice to have
- Supporting payment, fintech or banking systems, especially POS, switching or e-commerce platforms
- ISO 8583 host interfaces, terminal connectivity or HSM connectivity
- YugabyteDB, including multi-datacentre or active-active deployments
- Helm, kubectl-based troubleshooting, and GitOps or CI/CD-driven deployments
- Prometheus, Grafana, Loki or an ELK stack
- Wireshark or tcpdump
- Bash, PowerShell or Python scripting for diagnostics and automation
- Work in a PCI DSS certified environment
- CKA or CKAD, RHCSA or RHCE, Microsoft SQL Server or Windows Server certifications, CCNA or ITIL Foundation
- Arabic
How we work
- Systematic diagnosis: you work from evidence to root cause across layers and confirm the cause before you fix
- Composure under pressure: calm, methodical and communicative during major incidents
- Discipline: you follow deployment, change and rollback procedures exactly, especially when time is short
- Ownership: you stay with an issue until its recurrence is prevented, beyond just restoring service
- Breadth with depth: comfortable across OS, network, database and Kubernetes, and able to go deep in any of them
- Knowledge sharing: runbooks and coaching that make the whole support chain stronger
Where and when
Full-time, on-site at our Riyadh office. You'll be part of an on-call rotation supporting the 24/7 support centre, and will work outside standard hours for incidents, deployments and planned maintenance. Occasional travel within Saudi Arabia to datacentre, partner or customer sites may be required.
-