> checking production environment ... OK
> checking SLA compliance ... OK
> checking on-call readiness ... OK
> loading engineer profile ... DONE

Udikha Patsa Technical / Application Support Engineer — Production Support & Incident Management

Seven years keeping business-critical trading, banking, and healthcare applications online — incident management, root cause analysis, and 24×7 production support across ServiceNow, Grafana, Kibana, JIRA, and the FIX protocol. ITIL 4 and AZ-900 certified.

7+
Years in production support
3
Industries: trading · banking · health
0
SLA breaches, current role
ITIL 4
Foundation certified
Ticket desk

Four incidents, closed out

Real production issues from the current and last three roles — the kind that page you at 2am.

TKT-001 · Wio Invest & Securities

Trade execution & order routing was failing across ADX/DFM

SituationFIX session drops, execution rejects, and GTN OMS routing anomalies were destabilizing live trades across local and international markets.
SolutionTraced FIX tags, session logs, and order-status data across MongoDB, SQL Server, PostgreSQL, and Oracle; partnered with the OMS vendor and dev team to fix routing via BHM.
ResultSole L2 owner of 24×7 support for US/UAE equities, ETFs, structured products, and crypto trading since Nov 2023.
Resolved
TKT-002 · Payments & Banking Ops

Stripe & Mambu integrations were breaking transactions and invoices

SituationPayment failures, refund errors, and VAT/TRN mismatches in ExpressTXR were slipping into production.
SolutionDiagnosed via Postman/Curl, fixed the failing integrations, corrected the VAT/TRN discrepancies, and validated WPS Pro releases through Mixpanel and Airship before launch.
ResultClean payment flows, accurate invoicing, and validation built into every release cycle.
Resolved
TKT-003 · Dubai Health Authority

HIS/EHR/LIS platforms were breaching SLA, repeatedly

SituationAlhosn, Hasana, and Salama were missing uptime SLAs with slow escalation to onsite tech leads.
SolutionRebuilt the Ivanti incident workflow, added a daily SLA-breach review, and standardized outage reporting backed by Kibana log analysis.
ResultResponse times driven to zero recorded SLA breaches.
Resolved
TKT-004 · Optum Global Solutions

Support intake was fragmented across phone and email

SituationNo unified system existed to log, assign, or track requests across HL7/FHIR and PACS-integrated healthcare systems.
SolutionDesigned and shipped a new tracking system for coordinating support end-to-end.
ResultFaster response times and less downtime, without breaking HIPAA/GDPR compliance.
Resolved
Diagnostic tool

The 5 Whys machine

Click through the actual root-cause chain behind TKT-001 — the trading outage above.

Case file: TKT-001 · trade-routing.log
01
Why did trade execution fail?
Orders weren't routing correctly to ADX and DFM.
02
Why weren't orders routing correctly?
GTN OMS was sending malformed order messages via BHM.
03
Why were the messages malformed?
FIX sessions were dropping mid-transmission, corrupting the message sequence.
04
Why were FIX sessions dropping?
Session-connectivity checks weren't catching intermittent timeouts before they cascaded.
05
Why weren't timeouts caught earlier? (root cause)
No proactive session-health monitoring existed — alerts only fired after an order had already failed.
Fix applied → Built proactive FIX session monitoring into the Grafana/Kibana stack, so connectivity issues surface before an order is ever placed.
1 / 5
Ops room

Before this role, after this role

Drag to compare a manual, spreadsheet-driven tracker against the monitoring stack now running in production.

Before — manual tracker
TKT-2291SLA overdue
TKT-2292unassigned
TKT-2293no update 3d
TKT-2294SLA overdue
TKT-2295escalated by email
After — Grafana / Kibana / JIRA
100%
SLA compliance
0
Open incidents
Auto
Alert → JIRA ticket
24×7
On-call coverage

Illustrative mockup based on the real monitoring stack (Grafana, Kibana, JIRA) used in production — not an actual screenshot.

Health checkup

Run a production checkup

This is the actual daily routine — health checks across banking and trading applications, condensed into one click.

Wio Invest & Securities — order routing (ADX/DFM)
GTN OMS — FIX session connectivity
Payment gateway — Stripe / Mambu
Database layer — MongoDB / PostgreSQL / Oracle
Monitoring stack — Grafana / Kibana / Datadog alerts
On-call rotation — coverage confirmed
My SLA

The number that matters most

Not a satisfaction score — an SLA record, tracked since day one on the trading desk.

100% SLA compliance
0
SLA breaches, current role
7+
Years production uptime
24×7
On-call coverage
Standard operating procedure

SOP-0001

Doc: SOP-0001Owner: udikha_patsaRev: 1.0

Deploying a Production Support Engineer

Purpose

Eliminate production blind spots across trading, banking, and healthcare systems before they turn into incidents.

Scope

Applies to any team running 24×7 mission-critical applications without a dedicated L2 production support owner.

Procedure

  1. Onboard Udikha to ServiceNow / Ivanti / JIRA and grant read access to Grafana, Kibana, and Datadog.
  2. Connect the on-call rotation and existing SLA dashboards.
  3. Route the current backlog of unresolved incidents for triage and root cause analysis.
  4. Watch SLA breaches trend toward zero within the first cycle.
  5. Redeploy previously firefighting engineers back to feature work.
Effective: immediately upon offer acceptance. Approved by: udikha_patsa · 7 yrs production support.
Simulator

SIM_0314 — six minutes to market open

A real category of 3am decision. Pick a response and see what happens.

The FIX session to ADX drops mid-reconnect. Orders queued for the open are stuck in a pending state. Six minutes until the market opens. What do you do?

Endorsements

Reserved for the people who'd vouch for me

I'd rather leave this empty than put words in someone's mouth — these slots are open for real quotes from managers, clients, or teammates.

Engineering / Reporting Manager
— quote pending —
Cross-functional teammate (Dev / QA / SRE)
— quote pending —
Client / stakeholder
— quote pending —

Reach out on LinkedIn if you've worked with me and want to add one.

Contact

Get in touch