OCI Principal TPM — Technical Round
Data Center Infrastructure Delivery · Interviewer: Louis Chia
–
Days
–
Hours
–
Mins
–
Secs
📅 Tue, 29 Sep 2026 · 3:00 PM SGT · 60 min · Zoom · on camera · badge ready
Round 1 ✓
Recruiter pre-screen
Round 2 — now
Technical (elimination)
Ahead
Further rounds (4–5 total)

🎯How to answer (every time)

(1) direct answer in one line → (2) the mechanism / how you ran it → (3) one concrete example → (4) stop.
Louis is a career TPM (ex-AWS). lead with the mechanism — cadence, who owned mitigations, how you escalated — not just the headline. Frame in AWS language: working backwards, operational readiness review, correction-of-errors, mechanisms over intentions. At the edge of your knowledge: say what you know, how you'd close the gap, who you'd pull in. Never bluff — a Principal TPM orchestrates experts, they don't pretend to be one.
AI-use rule: Oracle permits AI in the interview only if openly declared and the interviewer opens it. No covert AI, no second screen, no whisper tool. Everything here has to be in your head — that's the whole point of rehearsing it.

🎬Your opener (first 90 seconds)

Louis will open with "give me the 90-second version." This is the one answer you can fully script. present tense first, then the arc, then why OCI, then pause and let him pick a thread.

I'm a data-centre delivery specialist, 15+ years across APAC. Right now I run independent end-to-end delivery for hyperscale builds in the region, Johor and Jakarta, taking sites from design and construction through commissioning to operational handover.

Before that I was the first data-centre operations hire at ByteDance in Singapore, where I built the delivery and operations function from zero and scaled it across APAC. I owned new-region builds from first steel to live production, 50-plus MW across 100-plus data halls in Singapore, and an 80 MW campus in Malaysia, up to 20 MW under concurrent build. Earlier I was in AWS Data Center Engineering Operations across Singapore, Tokyo and Seoul, so I've lived the M&E and mission-critical side hands-on.

What draws me to OCI is that this is exactly the high-density AI infrastructure delivery I do best, GPU halls, liquid cooling, tight timelines, and it's on the ground where I already work. I've spent the past year deepening my OCI knowledge through the Foundations, AI and Generative AI certifications.

So in short: I turn AI hardware into operating buildings, on time, on budget and to standard.
Delivery: ~75–90 seconds, unhurried. land the numbers cleanly, end on "why OCI", then stop and let Louis lead. don't narrate your whole CV.

⏱️The one-hour shape

An hour is long enough to go deep. Fewer questions than the screen, but heavy follow-up drilling on two or three areas. A considered 2–3 minute answer with a real example beats a fast shallow one.

TimeWhat's happening
0–5 minIntros; the interviewer frames their role and the session
5–45 minTechnical + behavioral deep-dive. they pick a few areas and keep asking "and then what / why / how" until they reach the edge of your knowledge
45–55 minYour questions
55–60 minWrap and next steps
They will follow up. Have depth behind every headline — likely deep-dive targets: delivery lifecycle, cooling / M&E, risk, and one or two STAR stories. Being drilled hard is normal, often a good sign. hit your limit, say so and explain how you'd close the gap.

🔒Lock these before you speak

🧭Who you're facing — Louis Chia

1. Cooling is the topic to nail. it's his top skill. go deep, and lead with your liquid-cooling SAT-failure RCA.
2. Program rigor over name-dropping. he's a career TPM — he'll ask how you ran it, not just that you did.
3. Complementary AWS DNA. you were AWS DCEO (facilities); he was AWS edge/network. you've delivered the plant he now program-manages. don't be intimidated.
4. Honest at the fabric boundary. your lane is the physical network layer (cabling, tray, circuits), not Infiniband fabric design. say so cleanly.

Source: his LinkedIn profile, confirmed.

🗣️Language to weave in — and which framework it's from

Three or four per answer, naturally. Never use a term you can't unpack. For Louis the AWS terms are gold — but never call an Amazon term "PMI" if a PMP-holder is in the room.

✅ PMI / PMBOK — safe to call PMI
Critical path (CPM) · dependencies (mandatory/discretionary/external/internal) · RACI / Responsibility Assignment Matrix · risk register (probability, impact, proximity) · risk responses (avoid, transfer, mitigate, accept, escalate) · contingency vs fallback plan · lessons learned (not "postmortem") · phase gate
⭐ AWS / Amazon — Louis's native tongue, your highest value
ORR (operational readiness review = your commissioning/handover gates) · COE (correction of errors = blameless RCA + 5 Whys = your SAT story) · mechanism (a process that enforces the fix = your water-quality gate) · single-threaded owner · working backwards · input vs output metrics · two-way / one-way door · Leadership Principles (Ownership, Dive Deep, Deliver Results, Bias for Action, Highest Standards, Earn Trust, Have Backbone). Recognise: bar raiser · Well-Architected / Operational Excellence · andon cord · Weekly Business Review · 6-pager narrative · tenets · Day 1.
⚠️ Common practice / other frameworks — use, but don't call PMI
RAID log (popular, not a PMBOK artifact; "D" = Dependencies or Decisions) · IMS (from EVM/defense, not PMBOK) · Definition of Done (Scrum) · blameless postmortem (DevOps; PMI = lessons learned) · cutover/go-live (ITIL) · DACI is Atlassian's, RACI is PMI's — don't mix them
Attribution discipline: the four AWS terms that do the most work, each tied to a real story, are ORR, COE, mechanism, single-threaded owner. Lead with those. If asked "is that PMI?", say honestly which bucket it's from.

❄️AI hardware & cooling — Louis's world

Speak it fluently and tie it to delivery + thermal, which you own. Figures are planning-level and evolve.

GB200 NVL72
A rack-scale system, not a single server. the GB200 "superchip" pairs 2 Blackwell GPUs with 1 Grace CPU; the NVL72 packs 72 Blackwell GPUs + 36 Grace CPUs across 18 compute trays and 9 NVLink-switch trays. 5th-gen NVLink joins all 72 GPUs into one NVLink domain — they act as a single giant GPU with pooled memory, which is what makes it so good for training huge models. The copper NVLink spine on the back is that all-to-all fabric. ~120–132 kW/rack, 100% liquid-cooled (direct-to-chip), no air-cooled option. Delivery lens: one rack carries the power and cooling load of a small legacy hall — it drives busway sizing, CDU, facility water and N+1 redundancy. Nvidia GB200 NVL72 →
GB300 NVL72
The Blackwell Ultra generation, the current top of the Blackwell line. Same NVL72 rack architecture as GB200 (72 GPUs + 36 Grace, one NVLink domain, copper spine), but upgraded B300 silicon: more FP4 compute and more HBM3e memory (~288 GB/GPU) for bigger models and longer context. That performance costs power, so the rack runs ~135–140 kW+ and leans even harder on the liquid loop. It reuses the same NVL72 power/cooling footprint, so for delivery it's a higher thermal envelope on familiar infrastructure. Visually indistinguishable from a GB200. Nvidia GB300 NVL72 →
AMD Instinct MI355X
AMD's answer to Blackwell (CDNA 4). Unlike the rack-scale NVL72, AMD's building block is a server: eight GPUs on an OAM baseboard (UBB) in a 4U liquid-cooled (or 8U air) chassis, linked by Infinity Fabric inside the node and a scale-out network between nodes. ~1.4 kW/GPU, liquid-cooled, 288 GB HBM3e. Matters to you because OCI runs AMD Instinct (MI300X, now MI355X) as a real alternative to Nvidia, so an OCI hall may run both, and both are high-density liquid-cooled. The Register: MI355X →
Infiniband (Quantum-X800)
The scale-out fabric that connects racks into a cluster. keep the two networks straight: NVLink is scale-up (GPU-to-GPU inside a rack), Infiniband is scale-out (between racks). Quantum-X800 is the current top: 800 Gb/s XDR, RDMA for ultra-low latency, in a rail-optimised non-blocking topology; switches are the Q3400 (4U, 144 ports) and Q3200 (2U). The Ethernet alternative is RoCE / Nvidia Spectrum-X. For you it means an explosion of high-count fibre and cable tray, the MMR/MDA layout and clean labeling — plan it as a critical-path dependency, because no fabric means no server bring-up.

Rail-optimised non-blocking, in plain English: non-blocking = every GPU can talk to every other at full line rate at once (a fat-tree / Clos with no oversubscription), because AI training has all GPUs exchanging data every step. rail-optimised = each GPU position (a "rail", e.g. all the #1 GPUs) homes to its own dedicated leaf switch, so collective traffic crosses one hop instead of the whole tree. The topology design is the network architect's job; you deliver the fibre, tray, pathways and labeling it runs on. Nvidia Quantum-X800 →
Roadmap: what's next — Rubin
After Blackwell comes Vera Rubin (Rubin GPU + Vera CPU, the successor to Grace), announced at GTC 2026 and now entering production, with Rubin Ultra ~2027 and Feynman beyond. Blackwell (GB200 / GB300) is what deploys now; Rubin is the incoming wave. Talking point: "each generation pushes rack power and liquid-cooling density further — that's the delivery and thermal problem I focus on." Avoid quoting hard Rubin rack numbers, they're still settling. Nvidia: Rubin →
How direct-to-chip cooling works +
At 125–140 kW/rack air can't carry the heat, so it's captured at the die. The chain: plant (towers/dry coolers/chillers) → Facility Water System (primary) → CDU (heat exchanger + pumps + filtration, isolates the loops) → Technology Cooling System (clean secondary) → rack manifold → dry-break quick-disconnects → cold plates on the dies → back to CDU. The CDU keeps dirty facility water separate from the pristine coolant touching the chips.

Facts to speak in: cold plates capture ~70–80% of heat (higher with rear-door HX / liquid-cooled memory), so a hall is hybrid liquid + air. Warm-water supply ~32 °C (W32) to ~45 °C (W45), not chilled → free-cooling hours + better PUE. Note: your 25→27 °C / PUE 1.34 result was air cooling (same principle, air side), not liquid — keep the two separate. N+1 CDUs/pumps — loss of flow is thermal runaway in seconds. Water quality is THE failure mode (debris clogs microchannels → hotspots) = your SAT failure. Commissioning: flush → pressure-test → fill → bleed → verify flow/temp → leak-check → energise. Switches are still mostly air-cooled, so liquid rows sit beside air rows.
The cooling chain — two loops that never mix

Why liquid at all: at ~130 kW a rack makes too much heat to move with air. water carries roughly 4,000× more heat per litre than air, so you carry the heat in liquid. Simplest picture: like blood carrying heat from your hot core to your skin, then sweat shedding it to the air. a small clean loop hugs the chips, hands the heat to the building's water at the CDU (the two never touch, like a double boiler), and the building water dumps it outside through a cooling tower.

GPU die — the heat source too hot for air (up to ~1.4 kW each) Cold plate on the die liquid runs through, grabs the heat CLEAN LOOP CDU — the handoff heat crosses a wall; loops never mix BUILDING LOOP Cooling tower / dry cooler heat leaves to the outside air cool coolant back water back
What's actually in the loops
The chip-side loop is not tap water — purified (deionised) water, usually with ~25% propylene glycol, plus a corrosion inhibitor and a biocide, held to tight conductivity and particulate specs. The building loop is treated condenser water. Tap water's minerals, microbes and grit would clog the cold-plate microchannels — the water-quality failure that cost you a cooling SAT. (Single-phase: the water stays liquid; immersion cooling is the one that uses a non-conductive dielectric fluid, not water.)
CDU — two ways it rejects heat
Liquid-to-liquid: heat crosses into the building water, then towers / dry coolers (standard for GPU halls). Liquid-to-air: the CDU dumps heat straight to room air via a radiator — no facility water needed, but lower capacity. CDUs also come in-rack, in-row or standalone (that's placement, not method).
Hook: "Density jumped from ~10–15 kW to 125–140 kW/rack — that's the whole story, and it's why my liquid-cooling and PUE work matters now."
Hook: "I've lived the direct-to-chip failure mode: failed a cooling SAT on water quality, drove the RCA, made water-quality verification a standing acceptance gate."
Boundary: "My lane is the physical plant and cabling the fabric runs on; I partner with the network architects on the logical fabric design."

Sources: Pantheon (power), The Register (MI355X), Vertiv (direct-to-chip), Vertiv (CDU), DCD (water temps), ASHRAE water quality.

⭐Chirag's 6 questions → your story

Facts unchanged. framing is now a TPM's. gold terms are the ones Louis will recognize; LP = the leadership principle each lands.

A failure & how you handled it
During integrated systems testing, the direct-to-chip loop failed its minimum test. I took single-threaded ownership and ran it as a correction-of-errors: root cause was water quality, construction debris in the loop. Real output was the mechanism — water-quality verification is now a standing acceptance gate in the commissioning ORR, so it can't recur.
RCA → permanent gate
LP: Dive Deep · Ownership · Highest Standards · Earn Trust
A success
Treated it as a two-way-door experiment, reversible, so we moved fast. Set the input metric I controlled, supply temp 25 → 27 across two 4.5 MW air-cooled halls, measured the output metric, PUE. Landed 1.34, cited by IMDA at the Tropical DC Standard launch.
PUE 1.34 · cited by IMDA
LP: Deliver Results · Invent & Simplify · Earn Trust
Delivery was delayed & recovery
On the integrated master schedule, network move-in was gated by cable-tray completion, and it slipped. I called it a critical-path dependency, not a parallel task, escalated early with impact and options, re-sequenced around the real constraint, protected the cutover.
Go-live milestone held
LP: Bias for Action · Dive Deep · Deliver Results
Risk assessment / management
I run an active RAID log: each risk gets probability, impact, proximity, the milestone it threatens, an owner, a response with a fallback. Live halls beside active construction was high-proximity — whole-building view, ingress/access controls, sequencing, escalated the residual with the specific decision needed.
Zero contamination / safety incidents
LP: Ownership · Dive Deep · Earn Trust
Largest infrastructure project & your role
80 MW Malaysia campus, buildings A–D staged. Single-threaded owner from design reviews to handover. Working backwards from the readiness date, built the master schedule, kept long-lead equipment visible so procurement never became the critical path, stood up the team from zero with RACI-clear roles.
80 MW campus · team from zero
LP: Ownership · Think Big · Hire & Develop · Deliver Results
How many projects at once
Up to 20 MW concurrent. two buildings rising together + a third where the hall was ready before the shell. Ran them as one portfolio against a single master schedule, allocated scarce resources by critical path and safety, held a separate operational-readiness review before each go-live.
20 MW under concurrent build
LP: Deliver Results · Dive Deep · Frugality

📎Your CV-backed evidence

Anchor every answer to what's already on your CV, so nothing you say can look invented.

M&E, power & cooling (AWS)
generators, UPS, PDUs, chillers, CRAH/AHU, BMS, fire suppression; DCIM & environmental monitoring.
High-density / GPU (ByteDance)
containment, airflow, pressure and power-train optimization on high-density GPU and AI/ML compute.
Cooling / sustainability
Digital Realty trial, 25 → 27 °C, PUE 1.34, cited by IMDA (SS 697:2023).
Network & cabling
rack deployments & network physical-infra upgrades (ByteDance); structured cabling, labeling, network install (current); circuits, Ethernet, subsea cable, POP provisioning (Pacnet/Telstra).
Delivery governance
risk & change via Jira/ITSM, structured RCA, Lean/Six Sigma, MOP/SOP and runbooks.
Scale & speed
50+ MW / 100+ halls / 1,000+ racks (SG); 80 MW campus (MY); Reuters 100 → 600 racks in three months.
Leadership
built ByteDance's delivery & ops function from zero as first hire; teams of 21–40 technicians.

🔧Technical domains — full answers, tap to open

Each has the headline answer, the mechanism, a real example, and a "deeper if pushed" note for the follow-up drilling.

1 · End-to-end delivery lifecycle (likely opener) +
"Walk me from lease signing to handover to operations. Where does it go wrong?"
End-to-end means owning the outcome across the full lifecycle, not one workstream: requirements & site selection → lease → design & reviews → construction → procurement → install (rack, network, cabling) → commissioning & SAT → acceptance → operational handover. My job is to hold schedule, budget and quality across all of it and keep design, construction, colo/BTS partners, network and ops aligned. It goes wrong at the seams between workstreams, not inside one: cable tray not ready when network gear moves in; access constraints found late; construction handed over "complete" but not clean, which fails commissioning. So I pull risks forward into design reviews and manage the dependencies, not just the boxes.
Deeper if pushed: name the milestones — first steel, topping out, power energization, water-on, white-space ready, network turn-up, SAT/IST pass, customer/ops handover — and say which sit on the critical path.
2 · Commissioning & acceptance +
"Walk me through commissioning. What does 'operationally ready' mean?"
Commissioning is levelled — the standard is Levels L1–L5 ("Cx" just means commissioning, not a level). L1 factory acceptance (FAT) pre-ship → L2 delivery/component verification on site → L3 pre-functional/static (installed right, safe to energize) → L4 functional performance of each system → L5 integrated systems test (IST) — the whole plant under load, incl. failure and black-building / pull-the-plug tests to prove power and cooling ride through a utility loss. "Construction complete" is not "operationally ready." Ready means every system tested to L5, acceptance evidence signed, the punch list closed or owned, and SOPs/MOPs/EOPs and as-builts in operations' hands. I treat handover to ops as the real finish line, and stay involved from design so the site is operable, not just built.
Your hook: the direct-to-chip liquid-cooling SAT failure — failed the minimum test, ran it to root cause (water quality / debris), remediated, made water-quality verification a standing acceptance criterion. A perfect commissioning-discipline story.
3 · Power train & redundancy +
"Describe the power path from utility to rack, and how you design for redundancy."
Utility → MV switchgear → transformers → LV switchgear → UPS → PDUs/RPPs → busway or whips → rack PDU. Backup is standby generators with automatic transfer switches; dual-corded loads sit behind static transfer switches so a rack keeps power if one path drops. Redundancy in N terms: N (none), N+1 (one spare), 2N (mirrored path), 2N+1. The two that matter operationally: concurrent maintainability (service any component without dropping load, ≈ Uptime Tier III) and fault tolerance (survive an unplanned failure, Tier IV / 2N). I map the redundancy target to the customer and SLA, then make sure commissioning proves it under failure, not just on paper.
Deeper if pushed: UPS types — static double-conversion, flywheel/rotary, DRUPS; battery choice VRLA vs lithium-ion. And Uptime Tiers I–IV by name.
4 · Cooling (your strongest area) +
"How do you approach cooling? Experience with liquid cooling?"
Air cooling is chilled water to CRAHs (or DX CRACs), hot/cold-aisle containment, and economization where climate allows — harder in the tropics, which is exactly why the Singapore Tropical DC Standard work mattered. As density climbs with GPU/AI, air can't carry the heat, so you move to liquid: direct-to-chip cold plates, or immersion. Direct-to-chip runs a technology cooling loop through a CDU that isolates the IT loop from facility water. I've delivered direct-to-chip on high-density GPU halls and know its failure modes first-hand: leakage, connector integrity, and water quality. On one build we failed SAT — root cause was water quality, construction debris in the loop degrading performance at the chip. We flushed and re-treated, re-ran SAT, and I made water-quality verification part of acceptance. On efficiency (this is air cooling, separate from liquid) I ran the Digital Realty trial raising the operating temp 25→27 °C across two 4.5 MW air-cooled halls, landing PUE 1.34, cited by IMDA at the Tropical DC Standard (SS 697:2023) launch. keep this air-side credential distinct from the direct-to-chip liquid work.
Go deep here: name CDU, TCS loop, facility water, conductivity/particulate limits, leak detection. See the AI hardware & cooling section for the full loop chain.
5 · Capacity & density +
"How do you think about rack density and capacity planning?"
Density is kW/rack, and it drives everything — floor space, power topology, cooling choice. Legacy racks were 5–15 kW; AI/GPU racks now run 40 to 100+ kW, which is why liquid cooling is no longer optional. When I quote 50 MW supporting 1,000+ racks, that's high-density GPU — the megawatts are high relative to rack count because each rack draws far more than a conventional one. Capacity planning is matching committed IT load to power and cooling headroom, staging it building by building, and forecasting BOM and long-lead equipment so procurement doesn't become the critical path.
6 · Network physical infrastructure & cabling +
"How do you handle the network and cabling side of a build?"
The physical network layer gates everything — without connectivity you can't bring up a server — so I plan it as a critical-path dependency, not an afterthought. That means structured cabling (copper and fibre, incl. bulk fibre), the MMR/MDA layout, cable tray and pathways sized and installed ahead of equipment, clean labeling, and ISP/circuit contracts executed on time. I've lived the failure: network gear scheduled to move in while the cable tray wasn't ready, which would have stalled server power-up. I re-sequenced around the real constraint and drove the blocker to closure before it hit go-live. My earlier network-engineering background — circuits, Ethernet, POP provisioning — means I read this layer fluently.
Boundary with Louis: physical layer is your lane. Infiniband fabric topology/design is the network architects'. say so cleanly.
7 · Vendor, colocation / BTS & budget +
"How do you manage colo and vendor partners, and the budget?"
Most hyperscale delivery runs through partners — colocation, build-to-suit (BTS), low-voltage/fit-out vendors — so I lead through governance and clear accountability, not direct authority. I coordinate with design so construction and design are locked before contracting, drive clean handoffs with construction and colo/BTS partners, and manage tenant fit-out and LV integration through commissioning. My role is coordinating and governing their delivery, not the commercial sourcing: aligning vendors to the design, clean handoffs, managing fit-out and LV integration through commissioning, and holding them to schedule, scope and quality. The commercial side — selection, negotiation, contracts — sits with procurement; I feed them the requirements, capacity/equipment forecast and acceptance criteria. I catch scope gaps and missing acceptance criteria early — a gap found in design costs a fraction of the same gap found at handover.
8 · Risk management +
"How do you manage risk on a build? Give an example."
I run a risk register and work it actively. For each risk: probability, impact, proximity and which milestone or capacity commitment it threatens, then an owner and a defined response — mitigate where I can, with a fallback for anything material. I escalate early, while options are still open, and I bring leadership the impact, the options and the specific decision needed, not just the problem. Example: deploying live halls while construction continued next door — dust, vibration, access congestion threatening production-bound equipment. I assessed it whole-building, put controls around ingress and access, and factored the neighbouring risk into my deployment sequencing and safety plan. We deployed with no contamination or safety incident.
PMI framing if it fits: risk register / RAID, qualitative vs quantitative analysis, mitigate/avoid/transfer/accept, contingency vs fallback. Use lightly, only if you can defend each term.

🎲Curveballs & how to handle

🧯Handling your gaps (honest framing)

Three things you haven't done. pre-scripted, calm answers so you're never caught. honesty at the edge reads as senior.

Commercial ownership of vendors
"The commercial side — selection, negotiation, contracts — sat with procurement. I fed them the technical requirements, the capacity and equipment forecast and the acceptance criteria, and I governed the delivery relationship: aligning vendors to the design, clean handoffs, holding them to schedule, scope and quality."
Infiniband / fabric design
"My lane is the physical fabric — structured cabling, tray, circuits, MMR/MDA, labeling — which I plan as a critical-path dependency. The logical topology design (rail-optimised, non-blocking) I partner with the network architects on."
Engineering degree / background
Don't volunteer "not an engineer." Lead with: 15 years hands-on plus CDCMP / CDFOM / CDCS and PMP in progress. "I've delivered and commissioned the plant end to end, which is what the role needs."
"Ops, not pure project management"
"The work was end-to-end program delivery — milestones, dependencies, risk, vendors, handover. the title varied, the function didn't."

❓Questions to ask Louis

⛔Traps to avoid (glance before the call)

✅Rehearsal tracker

Final reminders: Answer → mechanism → real example → stop. Go deepest on cooling and delivery, that's your edge. Keep every number consistent (50 Singapore / 80 Malaysia / 20 concurrent). Say APAC. Pause after answering. Let Louis lead.

Prep hub for Bernard Tay · OCI Principal TPM technical round · full guide PDF · bernardtay.work