OCI Principal TPM — Technical Round
Data Center Infrastructure Delivery · Interviewer: Louis Chia
–
Days
–
Hours
–
Mins
–
Secs
📅 Tue, 29 Sep 2026 · 3:00 PM SGT · 60 min · Zoom · on camera · badge ready
Round 1 ✓
Recruiter pre-screen
Round 2 — now
Technical (elimination)
Ahead
Further rounds (4–5 total)

🎯How to answer (every time)

(1) direct answer in one line → (2) the mechanism / how you ran it → (3) one concrete example → (4) stop.
Louis is a career TPM (ex-AWS). lead with the mechanism — cadence, who owned mitigations, how you escalated — not just the headline. Frame in AWS language: working backwards, operational readiness review, correction-of-errors, mechanisms over intentions. At the edge of your knowledge: say what you know, how you'd close the gap, who you'd pull in. Never bluff — a Principal TPM orchestrates experts, they don't pretend to be one.
AI-use rule: Oracle permits AI in the interview only if openly declared and the interviewer opens it. No covert AI, no second screen, no whisper tool. Everything here has to be in your head — that's the whole point of rehearsing it.

⏱️The one-hour shape

An hour is long enough to go deep. Fewer questions than the screen, but heavy follow-up drilling on two or three areas. A considered 2–3 minute answer with a real example beats a fast shallow one.

TimeWhat's happening
0–5 minIntros; the interviewer frames their role and the session
5–45 minTechnical + behavioral deep-dive. they pick a few areas and keep asking "and then what / why / how" until they reach the edge of your knowledge
45–55 minYour questions
55–60 minWrap and next steps
They will follow up. Have depth behind every headline — likely deep-dive targets: delivery lifecycle, cooling / M&E, risk, and one or two STAR stories. Being drilled hard is normal, often a good sign. hit your limit, say so and explain how you'd close the gap.

🔒Lock these before you speak

🧭Who you're facing — Louis Chia

1. Cooling is the topic to nail. it's his top skill. go deep, and lead with your liquid-cooling SAT-failure RCA.
2. Program rigor over name-dropping. he's a career TPM — he'll ask how you ran it, not just that you did.
3. Complementary AWS DNA. you were AWS DCEO (facilities); he was AWS edge/network. you've delivered the plant he now program-manages. don't be intimidated.
4. Honest at the fabric boundary. your lane is the physical network layer (cabling, tray, circuits), not Infiniband fabric design. say so cleanly.

Source: his LinkedIn profile, confirmed.

🗣️TPM language to weave in

Louis is ex-AWS — speak his OS. Three or four terms per answer, naturally. never use one you can't unpack.

Schedule & delivery
integrated master schedule (IMS) · critical path · dependency (vs parallel task) · gating · working backwards · cutover / go-live · input metric vs output metric
Ownership & decisions
single-threaded owner (STO) · RACI · DACI (for a decision) · two-way-door vs one-way-door (reversible = decide fast)
Risk & governance
RAID log · risk register (probability / impact / proximity) · escalation path · blameless postmortem
Amazon mechanisms (highest value)
COE = correction of errors, your RCA in his words · ORR = operational readiness review, your commissioning/handover gates · Mechanism = a process that enforces the fix so it doesn't rely on anyone remembering
Leadership principles to demonstrate
Ownership · Dive Deep · Deliver Results · Bias for Action · Insist on the Highest Standards · Earn Trust · Have Backbone, Disagree & Commit

❄️AI hardware & cooling — Louis's world

Speak it fluently and tie it to delivery + thermal, which you own. Figures are planning-level and evolve.

GB200 NVL72
72 Blackwell GPUs + 36 Grace CPUs, one NVLink domain. ~125–130 kW/rack, 100% liquid-cooled. recognise it by the copper NVLink spine on the back.
GB300 NVL72
Blackwell Ultra refresh. same rack, more HBM3e, ~135–140 kW+. visually identical to GB200.
AMD MI355X
AMD's rival, ~1.4 kW/GPU liquid-cooled, 288 GB HBM3e, 8 GPUs (OAM) in 4U liquid / 8U air. OCI runs AMD too.
Infiniband (Quantum-X800)
Scale-out fabric between racks (NVLink = scale-up within). 800 Gb/s XDR, RDMA. for you = huge fibre + cable-tray density.
How direct-to-chip cooling works +
At 125–140 kW/rack air can't carry the heat, so it's captured at the die. The chain: plant (towers/dry coolers/chillers) → Facility Water System (primary) → CDU (heat exchanger + pumps + filtration, isolates the loops) → Technology Cooling System (clean secondary) → rack manifold → dry-break quick-disconnects → cold plates on the dies → back to CDU. The CDU keeps dirty facility water separate from the pristine coolant touching the chips.

Facts to speak in: cold plates capture ~70–80% of heat (higher with rear-door HX / liquid-cooled memory), so a hall is hybrid liquid + air. Warm-water supply ~32 °C (W32) to ~45 °C (W45), not chilled → free-cooling hours + better PUE (your 25→27 °C, PUE 1.34 story). N+1 CDUs/pumps — loss of flow is thermal runaway in seconds. Water quality is THE failure mode (debris clogs microchannels → hotspots) = your SAT failure. Commissioning: flush → pressure-test → fill → bleed → verify flow/temp → leak-check → energise. Switches are still mostly air-cooled, so liquid rows sit beside air rows.
Hook: "Density jumped from ~10–15 kW to 125–140 kW/rack — that's the whole story, and it's why my liquid-cooling and PUE work matters now."
Hook: "I've lived the direct-to-chip failure mode: failed a cooling SAT on water quality, drove the RCA, made water-quality verification a standing acceptance gate."
Boundary: "My lane is the physical plant and cabling the fabric runs on; I partner with the network architects on the logical fabric design."

Sources: Pantheon (power), The Register (MI355X), Vertiv (direct-to-chip), Vertiv (CDU), DCD (water temps), ASHRAE water quality.

⭐Chirag's 6 questions → your story

Facts unchanged. framing is now a TPM's. gold terms are the ones Louis will recognize; LP = the leadership principle each lands.

A failure & how you handled it
During integrated systems testing, the direct-to-chip loop failed its minimum test. I took single-threaded ownership and ran it as a correction-of-errors: root cause was water quality, construction debris in the loop. Real output was the mechanism — water-quality verification is now a standing acceptance gate in the commissioning ORR, so it can't recur.
RCA → permanent gate
LP: Dive Deep · Ownership · Highest Standards · Earn Trust
A success
Treated it as a two-way-door experiment, reversible, so we moved fast. Set the input metric I controlled, supply temp 25 → 27 across two 4.5 MW halls, measured the output metric, PUE. Landed 1.34, cited by IMDA at the Tropical DC Standard launch.
PUE 1.34 · cited by IMDA
LP: Deliver Results · Invent & Simplify · Earn Trust
Delivery was delayed & recovery
On the integrated master schedule, network move-in was gated by cable-tray completion, and it slipped. I called it a critical-path dependency, not a parallel task, escalated early with impact and options, re-sequenced around the real constraint, protected the cutover.
Go-live milestone held
LP: Bias for Action · Dive Deep · Deliver Results
Risk assessment / management
I run an active RAID log: each risk gets probability, impact, proximity, the milestone it threatens, an owner, a response with a fallback. Live halls beside active construction was high-proximity — whole-building view, ingress/access controls, sequencing, escalated the residual with the specific decision needed.
Zero contamination / safety incidents
LP: Ownership · Dive Deep · Earn Trust
Largest infrastructure project & your role
80 MW Malaysia campus, buildings A–D staged. Single-threaded owner from design reviews to handover. Working backwards from the readiness date, built the master schedule, managed long-lead procurement so the BOM never became the critical path, stood up the team from zero with RACI-clear roles.
80 MW campus · team from zero
LP: Ownership · Think Big · Hire & Develop · Deliver Results
How many projects at once
Up to 20 MW concurrent. two buildings rising together + a third where the hall was ready before the shell. Ran them as one portfolio against a single master schedule, allocated scarce resources by critical path and safety, held a separate operational-readiness review before each go-live.
20 MW under concurrent build
LP: Deliver Results · Dive Deep · Frugality

📎Your CV-backed evidence

Anchor every answer to what's already on your CV, so nothing you say can look invented.

M&E, power & cooling (AWS)
generators, UPS, PDUs, chillers, CRAH/AHU, BMS, fire suppression; DCIM & environmental monitoring.
High-density / GPU (ByteDance)
containment, airflow, pressure and power-train optimization on high-density GPU and AI/ML compute.
Cooling / sustainability
Digital Realty trial, 25 → 27 °C, PUE 1.34, cited by IMDA (SS 697:2023).
Network & cabling
rack deployments & network physical-infra upgrades (ByteDance); structured cabling, labeling, network install (current); circuits, Ethernet, subsea cable, POP provisioning (Pacnet/Telstra).
Delivery governance
risk & change via Jira/ITSM, structured RCA, Lean/Six Sigma, MOP/SOP and runbooks.
Scale & speed
50+ MW / 100+ halls / 1,000+ racks (SG); 80 MW campus (MY); Reuters 100 → 600 racks in three months.
Leadership
built ByteDance's delivery & ops function from zero as first hire; teams of 21–40 technicians.

🔧Technical domains — full answers, tap to open

Each has the headline answer, the mechanism, a real example, and a "deeper if pushed" note for the follow-up drilling.

1 · End-to-end delivery lifecycle (likely opener) +
"Walk me from lease signing to handover to operations. Where does it go wrong?"
End-to-end means owning the outcome across the full lifecycle, not one workstream: requirements & site selection → lease → design & reviews → construction → procurement → install (rack, network, cabling) → commissioning & SAT → acceptance → operational handover. My job is to hold schedule, budget and quality across all of it and keep design, construction, colo/BTS partners, network and ops aligned. It goes wrong at the seams between workstreams, not inside one: cable tray not ready when network gear moves in; access constraints found late; construction handed over "complete" but not clean, which fails commissioning. So I pull risks forward into design reviews and manage the dependencies, not just the boxes.
Deeper if pushed: name the milestones — first steel, topping out, power energization, water-on, white-space ready, network turn-up, SAT/IST pass, customer/ops handover — and say which sit on the critical path.
2 · Commissioning & acceptance +
"Walk me through commissioning. What does 'operationally ready' mean?"
Commissioning is levelled. L1 factory acceptance (FAT) pre-ship → L2 delivery/component verification on site → L3 pre-functional/static (installed right, safe to energize) → L4 functional performance of each system → L5 integrated systems test (IST) — the whole plant under load, incl. failure and black-building / pull-the-plug tests to prove power and cooling ride through a utility loss. "Construction complete" is not "operationally ready." Ready means every system tested to L5, acceptance evidence signed, the punch list closed or owned, and SOPs/MOPs/EOPs and as-builts in operations' hands. I treat handover to ops as the real finish line, and stay involved from design so the site is operable, not just built.
Your hook: the direct-to-chip liquid-cooling SAT failure — failed the minimum test, ran it to root cause (water quality / debris), remediated, made water-quality verification a standing acceptance criterion. A perfect commissioning-discipline story.
3 · Power train & redundancy +
"Describe the power path from utility to rack, and how you design for redundancy."
Utility → MV switchgear → transformers → LV switchgear → UPS → PDUs/RPPs → busway or whips → rack PDU. Backup is standby generators with automatic transfer switches; dual-corded loads sit behind static transfer switches so a rack keeps power if one path drops. Redundancy in N terms: N (none), N+1 (one spare), 2N (mirrored path), 2N+1. The two that matter operationally: concurrent maintainability (service any component without dropping load, ≈ Uptime Tier III) and fault tolerance (survive an unplanned failure, Tier IV / 2N). I map the redundancy target to the customer and SLA, then make sure commissioning proves it under failure, not just on paper.
Deeper if pushed: UPS types — static double-conversion, flywheel/rotary, DRUPS; battery choice VRLA vs lithium-ion. And Uptime Tiers I–IV by name.
4 · Cooling (your strongest area) +
"How do you approach cooling? Experience with liquid cooling?"
Air cooling is chilled water to CRAHs (or DX CRACs), hot/cold-aisle containment, and economization where climate allows — harder in the tropics, which is exactly why the Singapore Tropical DC Standard work mattered. As density climbs with GPU/AI, air can't carry the heat, so you move to liquid: direct-to-chip cold plates, or immersion. Direct-to-chip runs a technology cooling loop through a CDU that isolates the IT loop from facility water. I've delivered direct-to-chip on high-density GPU halls and know its failure modes first-hand: leakage, connector integrity, and water quality. On one build we failed SAT — root cause was water quality, construction debris in the loop degrading performance at the chip. We flushed and re-treated, re-ran SAT, and I made water-quality verification part of acceptance. On efficiency, I ran the Digital Realty trial raising supply temp 25→27 °C across two 4.5 MW halls, landing PUE 1.34, cited by IMDA.
Go deep here: name CDU, TCS loop, facility water, conductivity/particulate limits, leak detection. See the AI hardware & cooling section for the full loop chain.
5 · Capacity & density +
"How do you think about rack density and capacity planning?"
Density is kW/rack, and it drives everything — floor space, power topology, cooling choice. Legacy racks were 5–15 kW; AI/GPU racks now run 40 to 100+ kW, which is why liquid cooling is no longer optional. When I quote 50 MW supporting 1,000+ racks, that's high-density GPU — the megawatts are high relative to rack count because each rack draws far more than a conventional one. Capacity planning is matching committed IT load to power and cooling headroom, staging it building by building, and forecasting BOM and long-lead equipment so procurement doesn't become the critical path.
6 · Network physical infrastructure & cabling +
"How do you handle the network and cabling side of a build?"
The physical network layer gates everything — without connectivity you can't bring up a server — so I plan it as a critical-path dependency, not an afterthought. That means structured cabling (copper and fibre, incl. bulk fibre), the MMR/MDA layout, cable tray and pathways sized and installed ahead of equipment, clean labeling, and ISP/circuit contracts executed on time. I've lived the failure: network gear scheduled to move in while the cable tray wasn't ready, which would have stalled server power-up. I re-sequenced around the real constraint and drove the blocker to closure before it hit go-live. My earlier network-engineering background — circuits, Ethernet, POP provisioning — means I read this layer fluently.
Boundary with Louis: physical layer is your lane. Infiniband fabric topology/design is the network architects'. say so cleanly.
7 · Vendor, colocation / BTS & budget +
"How do you manage colo and vendor partners, and the budget?"
Most hyperscale delivery runs through partners — colocation, build-to-suit (BTS), low-voltage/fit-out vendors — so I lead through governance and clear accountability, not direct authority. I coordinate with design so construction and design are locked before contracting, drive clean handoffs with construction and colo/BTS partners, and manage tenant fit-out and LV integration through commissioning. Commercially I select, negotiate and govern vendors and subcontractors, own BOM and capacity forecasting, and hold projects to budget with cost control and change management. I catch scope gaps and missing acceptance criteria early — a gap found in design costs a fraction of the same gap found at handover.
8 · Risk management +
"How do you manage risk on a build? Give an example."
I run a risk register and work it actively. For each risk: probability, impact, proximity and which milestone or capacity commitment it threatens, then an owner and a defined response — mitigate where I can, with a fallback for anything material. I escalate early, while options are still open, and I bring leadership the impact, the options and the specific decision needed, not just the problem. Example: deploying live halls while construction continued next door — dust, vibration, access congestion threatening production-bound equipment. I assessed it whole-building, put controls around ingress and access, and factored the neighbouring risk into my deployment sequencing and safety plan. We deployed with no contamination or safety incident.
PMI framing if it fits: risk register / RAID, qualitative vs quantitative analysis, mitigate/avoid/transfer/accept, contingency vs fallback. Use lightly, only if you can defend each term.

🎲Curveballs & how to handle

❓Questions to ask Louis

✅Rehearsal tracker

Final reminders: Answer → mechanism → real example → stop. Go deepest on cooling and delivery, that's your edge. Keep every number consistent (50 Singapore / 80 Malaysia / 20 concurrent). Say APAC. Pause after answering. Let Louis lead.

Prep hub for Bernard Tay · OCI Principal TPM technical round · full guide PDF · bernardtay.work