Scan Bench
One page per interview topic, each with a 60-second answer, drawn schematics and timing diagrams, the traps interviewers like, and questions with model answers. Round 1 covers the first two focus areas; round 2 is scan PD, DFT timing and SDC. The deep dives go further on EDT, test-mode timing and SSN.
Focus areas
Scan architecture & clocking
Mux-D, shift/capture protocol, lockup latches, OCC, LOC vs LOS.
Scan insertion & compression
Scannability DRC, stitching and balancing, EDT, X masking.
Core wrapping
Wrapper cells, internal/external modes, graybox, retargeting, IEEE 1500.
Test point insertion
Random-pattern resistance, control/observe points, SCOAP, costs.
Low-power scan
Shift vs capture power, fill, Q-gating, ICG TE, UPF domains.
JTAG 1149.1
TAP state machine, registers, instructions, timing, BSR cell.
IJTAG 1687
TDRs, SIBs, ICL/PDL, retargeting.
IEEE 1500 wrapper
WSP/WSC, WIR/WBY/WBR, instructions, TAP/IJTAG control, CTL.
Fault models & ATPG
Models, D-alg/PODEM, fault classes, coverage formulas, AU debug.
GLS & SDF
Annotation, timing checks and X, serial vs parallel, debug method.
Post-silicon debug
Bring-up ladder, chain signatures, chain & logic diagnosis, shmoo.
EDT signals & behavior
Interface, edt_update and edt_clock cycle by cycle, masking, bypass, what breaks.
Test-mode timing
Scan-in from a GPIO pin, scan-out strobe, SE settling before capture, checks per mode.
Streaming Scan Network
Bus and hosts, packets, IJTAG setup, on-chip compare, and the switching-noise SSN.
Scripting drills
Flists, RTL vs netlist parsing, chain tracing, fail logs, with hints and partial solutions.
Scan PD, DFT timing & SDC
Reordering, SE, congestion, test modes, SDC per mode, convergence.
How to use this
- Rate every page first. Use the "My level" buttons at the top of each page (Daily, Solid, Rusty, Weak). The dots in the sidebar and on the cards above follow your rating. Ratings stay in this browser only.
- Study Weak and Rusty first. For each: read the 60-second answer aloud, study the figures until you can redraw them, then answer the questions before opening the model answers.
- Redraw from memory. The whiteboard figures that come up most: mux-D cell, lockup latch + its timing, OCC + capture waveform, LOC vs LOS, EDT architecture, wrapper cell, TAP state machine, SIB, chain-test signatures.
- Daily topics get a quick pass. For topics you use every day, read the traps and the "how to say it" boxes. The risk there isn't knowing the material; it's explaining your tool flow in general terms.
- Drill in chat. Mock interviews and scripting practice happen in the conversation with Claude (formats below).
Mock interview formats
Start any of these in the chat by typing the line in the box, or say it aloud in voice mode. Claude plays the interviewer, asks one question at a time, pushes with follow-ups, and scores at the end.
Rapid fire · 15 min
10–12 short questions across a topic set. One or two sentences each. Tests recall and precision.
mock rapid: scan + JTAGDeep dive · 30 min
One topic, whiteboard style: you describe the drawing, Claude drills into why, edge cases, and "what if".
mock deep: OCC and at-speedDebug scenario · 20 min
A GLS or silicon failure. You ask for data, Claude reveals it as the tester or simulator would. Scored on method, not just the answer.
mock debug: silicon chain failScripting · 30–45 min
A Python problem from the scripting drills page. You write pseudo-code or Python; Claude reviews for correctness, edge cases, complexity and clarity, then gives a follow-up twist.
mock script: flistScoring rubric (1–4 each)
| Dimension | 4 looks like |
|---|---|
| Correctness | No technical errors; terms used precisely |
| Depth | Explains why, names trade-offs, knows where it breaks |
| Structure | Answer first, then support; a method for open problems |
| Communication | Draws or describes a figure clearly; tool-agnostic terms with tool examples |
Scripting problem bank
The full set, with format reminders, hints and partial solutions, is on the Scripting drills page. Use it to prepare, then drill in chat without looking.
About the existing site content
This review was written independently of vlsi.tylerpeller.com so it can serve as a cross-check. To fact-check the site itself, connect its GitHub repository in the chat; each DFT page can then be compared against this material and corrections proposed file by file.
Scan architecture and clocking
How a scan cell works, how shift and capture are sequenced, and how clocks are controlled for stuck-at and at-speed test: lockup latches, OCCs, LOC and LOS.
60-second answer
Scan turns a sequential test problem into a combinational one. Every flop gets a mux (SE selects between functional D and scan-in), and the flops are stitched into shift registers. In shift mode the tester loads state and unloads responses serially at a slow, safe frequency. In capture mode SE is low and one or more functional clock pulses capture the logic's response. For stuck-at, one slow capture pulse is enough. For transition faults we need a launch and a capture exactly one functional period apart, which in practice comes from an on-chip clock controller that gates PLL pulses. Clock-domain crossings inside a chain get lockup latches, and all of it has to be timed in separate shift and capture modes.
The scan cell
Mux-D is the industry default: one standard cell (e.g. SDFF) with pins D, SI, SE, CK, Q, sometimes a separate SO. It costs a mux delay on the functional D path (the scan mux sits in the setup path) and a little area. Two other styles you should be able to name:
- LSSD (level-sensitive scan design, IBM): two latches with non-overlapping clocks. Master latch takes either system data (clock C) or scan data (clock A); slave latch clocked by B. Race-free by construction and robust to hold problems, but needs extra clock routing. Rare in modern SoCs.
- Clocked scan: separate scan clock selects the scan input instead of a mux. Avoids the mux delay in the functional path.
Chains and the shift/capture protocol
scan_in pin to scan_out. SE (orange) and the clock (blue) are shared by every cell in the chain. With SE=1 the tester controls every flop state; with SE=0 one capture pulse turns the circuit into a pure combinational test problem.Everything a chain touches is shared: SE and the shift clock go to every cell, so a single defect in the SE tree or the clock tree can break every chain at once. That is the first thing you rule out in silicon bring-up.
Stuck-at pattern sequence
- Load (SE=1): shift L cycles; the chain now holds the pattern. The previous response comes out on scan_out at the same time.
- Force PIs, then measure POs (SE=0). PO strobe is before the capture pulse, because capture changes flop state.
- Capture: one clock pulse; every flop samples its functional D.
- Unload overlapped with the next load.
test time ≈ test cycles / fshift
Example: 10,000 patterns, longest chain 2,000 cells, 50 MHz shift → ~20.0 M cycles ≈ 0.4 s. Halve the longest chain and test time roughly halves. That's the case for balanced chains and compression. Only the longest chain matters, because all chains shift in parallel.
Mixed clocks in a chain: lockup latches and edge ordering
Chains often cross clock domains or clock-tree branches with large skew between them. In shift mode all of these are driven from one shift clock, but the insertion delays differ. If the capture flop's clock is later than the launch flop's clock by more than the Q→SI delay, the capture flop takes the new bit: a hold violation in shift. The chain loses a bit.
- Rule: for positive-edge flops, the lockup latch is a negative-level (active-low) latch clocked by the launching domain's clock. Tools insert them automatically at domain boundaries in a chain (
set_scan_configuration -add_lockupstyle options, or Tessent's lockup cell insertion when stitching across clocks). - A lockup latch does not add a cell to the chain; it retimes the transfer by half a cycle. A lockup flop would add a bit, which also works but changes the chain length.
- Mixed edges: put negative-edge flops before positive-edge flops in a chain (or keep them on separate chains). A pos→neg hop inside one cycle lets a bit move through two cells in one shift cycle; neg→pos is a clean half-cycle transfer.
- If the launch clock is the late one, you have a setup issue instead, and it's harmless at slow shift frequency.
Clock control: why an OCC
For at-speed test the capture pulses have to be at functional frequency (hundreds of MHz to GHz). Testers can't drive that cleanly through package pins, and the real functional clock comes from an on-chip PLL anyway. An on-chip clock controller sits between the PLL output and each clock domain's root. It chooses between the slow shift clock and a precise burst of PLL pulses.
- Clock-chain bits
- A short shift register inside the OCC, stitched into a scan chain. ATPG loads it per pattern, so each pattern chooses how many pulses (e.g. 1 for stuck-at, 2 for LOC) and which domains pulse. This is what lets the ATPG tool model the OCC as a programmable clock source (Tessent handles this through clock control definitions for the OCC).
- SE synchronizer
- SE comes from the tester, asynchronous to the PLL. It's synchronized into the fast domain so the pulse burst starts cleanly and at a fixed cycle count.
- Glitch-free gating
- Gating uses a latch-based ICG, and the slow/fast switch is a glitch-free clock mux. Any glitch on a clock root corrupts the whole domain.
- Slow capture
- Stuck-at patterns can capture with a single shift-clock pulse routed through the OCC (the "slow" path), useful for debug and for domains with no PLL.
- Inter-domain at-speed
- Launch in domain A, capture in domain B, at speed, is only meaningful if the two OCCs are synchronous and aligned. Otherwise ATPG pulses one domain per pattern (or compatible groups) and masks cross-domain paths.
At-speed launch: LOC vs LOS
| LOC (broadside) | LOS (skewed-load) | |
|---|---|---|
| Launch | First functional capture pulse | Last shift pulse |
| V2 | Circuit response to V1 (functionally reachable) | V1 shifted by one cell (any value ATPG wants) |
| SE timing | Slow; SE settles before launch | At-speed: must fall between launch and capture |
| Coverage / patterns | Lower coverage per pattern, more patterns, slower ATPG | Higher coverage, fewer patterns |
| Overtest risk | Low: tests reachable transitions | Higher: can launch non-functional transitions and fail good parts |
| Hardware | Just OCC/clock control | Pipelined SE or at-speed SE tree, per-domain |
| Industry use | LOC is the default in most flows. LOS is used selectively, sometimes as a top-off. Hybrid schemes exist. | |
Interview questions
Why does scan make ATPG tractable? basic
Without scan, detecting a fault means finding an input sequence that drives the state machine into the right state and then propagates the effect to a pin. That's sequential ATPG and it's very expensive. With full scan every flop is directly controllable (load) and observable (unload), so each flop output becomes a pseudo-PI and each flop input a pseudo-PO. ATPG only has to solve the combinational logic between them.
Where exactly do you put a lockup latch and what clocks it? core
Between the last cell of domain A and the first cell of domain B in the chain, clocked by domain A's clock (the launch side). For posedge flops it's a latch that's transparent when that clock is low. It holds the old value through the high phase, which gives half a period of hold margin to absorb skew between the domains. It does not add a scan bit.
A chain has a hold violation in shift only at one clock-domain crossing. What are your fixes? core
Add a lockup latch if it's missing (or check it's clocked by the right domain). Otherwise add hold buffering on the Q→SI path, reorder the chain so the crossing sits between closer clock-tree leaves, or use the flop's dedicated SO pin if there is one. Also check the shift-mode timing constraints: hold on shift paths must be analyzed with shift-mode case analysis and real clock latencies.
Why can't the ATE just supply the at-speed clock? core
GHz clocks through package pins and probe cards suffer from signal-integrity problems, tester channel limits and cost. More importantly, the functional clock tree is designed around the PLL output. Testing with an external clock doesn't exercise the real clock path and has different jitter/duty cycle. The OCC reuses the PLL and only needs slow control signals from the tester.
How does ATPG know how many pulses the OCC will deliver? core
The OCC's control bits (clock chain) are scan cells. The tool is given a model of the OCC: which bits enable which pulse in the capture window. It then treats the bits as part of the pattern. It chooses 1 pulse for a stuck-at pattern, 2 for a LOC transition pattern, sometimes more for sequential depth, and sets them in the load data.
What's the risk of LOS for yield? deep
LOS can create V1→V2 transitions that are functionally impossible, e.g. exercising a path that's a false path or a multicycle path in functional mode. A good chip fails at speed on a path that never matters. You mitigate with timing exceptions fed to ATPG (false/multicycle paths masked), or by using LOC as the main method.
Why must PO be strobed before the capture pulse in a stuck-at pattern? basic
The expected PO values are computed from the loaded state + forced PIs. After the capture pulse the flops take new values, and any PO that depends combinationally on flop outputs changes. Standard order: force PI → measure PO → pulse capture clock.
What is a "dead cycle" around SE, and why add one? core
An extra tester cycle with no clock pulse after SE toggles, so the high-fanout SE net settles everywhere before the next active edge. It lets SE be timed as a multicycle path, which relaxes its buffering. The OCC's synchronizer delay provides a similar settling window on the capture side.
Your chain includes both posedge and negedge flops. What goes where? core
Negedge flops first (closer to scan-in), posedge flops after. Within one shift cycle, the posedge flops shift at the rising edge and then the negedge flops at the falling edge. If a negedge flop follows a posedge one, it grabs the bit the posedge flop just took, so two cells shift in one cycle. Alternatives: separate chains per edge, or a lockup element at the boundary.
Further reading: EDN, launch-off-shift at-speed test · Siemens, navigating Tessent OCCs · VLSI Tutorials, on-chip clock controller
Scan insertion and compression
What has to be true of a netlist before it can be scanned, how chains are built and balanced, and how compression cuts test data and time by one to two orders of magnitude.
60-second answer
Insertion replaces flops with scan flops and stitches them into chains, but most of the engineering is making the design scannable. Every scan flop's clock and async set/reset must be controllable from the tester, there must be no X sources the tool can't handle, and memories and black boxes need to be isolated or bypassed. Then chains are built per clock domain, edge and power domain, and balanced. Compression puts a decompressor/compactor pair between a few tester channels and hundreds of short internal chains. That works because ATPG patterns are mostly don't-cares, so a small stream of channel data can encode the few care bits.
Where insertion sits
Scannability rules (the stuff that breaks ATPG)
| Problem | Why it breaks scan | Typical fix |
|---|---|---|
| Generated / gated / muxed clocks | Flop clock not controllable from a test clock source, so shift fails | Test-mode clock mux to OCC/test clock; ICG TE pin tied to SE |
| Async set/reset from logic | Reset can fire during shift and wipe the chain | Gate or mux the reset in test mode; hold it off during shift; ATPG can pulse it in capture |
| Latches | Not scannable in a mux-D flow | Make transparent in test mode, or treat as sequential elements ATPG can model |
| Combinational loops | Oscillation, unknown values | Break with a test-mode gate or scan flop |
| Tristate buses | Contention (two drivers on) or float (none) | One-hot enable decode in test, bus keepers; ATPG contention checks |
| Memories, analog, hard macros | X sources; logic around them untestable (shadow logic) | Bypass/observe logic around the macro, or macro test mode; model as black box and mask |
| Non-scan flops | Uncontrollable state, X at capture | Scan them, reset them in test, or accept as X source and mask |
| Pad/IO control | Bidirectional pins toggle direction during test | Force direction in test mode; boundary scan controls pads |
Stitching and balancing
Chain construction rules, roughly in order of priority:
- Clock domain and edge: keep one domain per chain where possible. Where not, order and add lockup latches (see scan clocking). Negedge before posedge.
- Power domain: a chain should not cross a switchable domain, or the chain breaks when that domain is off (and needs isolation/level shifters when it's on).
- Hierarchy / partition: wrapper chains and internal chains per core; keep physical partitions together.
- Balance: equal length within each compression unit or channel group.
- Physical: reorder after placement using the scan DEF, within those constraints (see round 2).
Compression
A pattern typically has only 1–5% care bits. Compression exploits that: the tester supplies a small compressed stream, and on-chip hardware expands it. Test data volume and test time drop by roughly the chain-to-channel ratio, until either encoding capacity or X density limits it.
- Decompressor
- Continuous-flow linear decompressor: ring generator (a compact LFSR) + phase shifter. Each internal chain bit is a linear (XOR) function of the channel bits injected so far. ATPG solves a system of linear equations over GF(2) per pattern to find channel data that produces the care bits. If the care bits don't fit, that fault moves to a different pattern.
- Compactor
- Spatial XOR trees combine many chain outputs per channel each shift cycle. Aliasing (two errors canceling) is possible but rare and accounted for by fault simulation.
- Masking
- Per-pattern (or per-shift) mask bits, delivered with the compressed data, block chains carrying X. Masking also reduces observability, so heavy X sources cost coverage and patterns.
- Ratio
- Compression ratio ≈ internal chains / scan channels, adjusted by pattern inflation. Past a point, more chains means more patterns (encoding limits) and more routing congestion around the decompressor and compactor, so the effective gain flattens.
- Bypass
- A bypass mode concatenates internal chains into long uncompressed chains. It's the first thing you use in silicon debug because failures map directly to cells.
- Low-pin-count
- Channel counts can be tiny (a few pins) for package test or for hierarchical cores sharing top-level pins through broadcast or channel sharing.
e.g. 2 M cells, 8 channels: no compression, 8 chains of 250 k cells. With 800 internal chains: 2.5 k cells/chain, ~100× shorter shift, typically 1.5–3× more patterns → ~30–60× net.
X handling
Where X comes from, and what to do:
- Uninitialized state (non-scan flops, memories, register files): initialize in test setup, use memory bypass, or constrain.
- Black boxes / analog: model outputs as X, add observe/control logic around them.
- Timing exceptions in at-speed capture: false and multicycle paths can't be trusted at speed, so ATPG masks their endpoints (captures X). That's why SDC exceptions get imported into ATPG.
- Cross-domain captures with asynchronous clocks: mask or don't pulse both domains in one pattern.
- Bus contention / floating: ATPG resolves to X.
Interview questions
Why does compression work at all? basic
Deterministic ATPG patterns specify only a small fraction of bits (care bits); the rest are don't-cares. A linear decompressor can produce any small set of specified bits from much less input data, as long as the linear equations are solvable. The don't-care bits come out as whatever the decompressor happens to generate.
What limits the compression ratio? core
(1) Encoding capacity: care bits per shift cycle can't exceed roughly what the channels inject, so dense patterns fail to encode and faults move to extra patterns. (2) X density and masking reduce observability. (3) Physical: hundreds of chains fanning out of a decompressor and into a compactor cause routing congestion. (4) Aliasing is minor. In practice pattern count rises as the ratio rises, so the net benefit saturates.
ATPG shows low coverage on logic around an SRAM. What's happening and what do you do? core
That's shadow logic: the memory's outputs are X to ATPG and its inputs aren't observed, so logic feeding and fed by the RAM is uncontrollable or unobservable. Options: memory bypass mode (a mux from data-in to data-out in test mode), observation flops on the RAM inputs, ATPG with a RAM model (RAM sequential patterns), or accept and cover the memory with MBIST.
How do you decide how many scan chains / channels? core
Start from test time and pin budget. Channels are limited by available test pins (and the tester's channel memory). Chain length is set by the target: total cells / internal chains. Then check routing congestion around the EDT logic, shift power (more parallel chains are fine, total toggling is the same), and whether hierarchical cores share channels. Iterate with pattern count estimates.
Why a compression bypass mode? basic
For debug and diagnosis. In bypass, a tester fail maps directly to (chain, cell) without having to invert the compactor. Also for chain integrity tests before trusting the EDT logic, and as a fallback if the compression logic has a defect.
Where do you place chain-to-domain crossings and why? core
As few as possible, grouped, with lockup latches at each crossing. Many tools order chains by clock domain and edge so there's at most one crossing per group, and put the lockup latch right at it. Crossings are the most common source of shift hold issues after CTS.
Reset is asynchronous and generated by a reset synchronizer. What DRC do you expect and how do you fix it? core
An "uncontrollable set/reset" violation: the reset output of the synchronizer isn't directly controllable from a PI during shift. Fix: in test mode, mux the flops' reset to a top-level pin (or gate it with a test signal so it's inactive during shift). Also scan the synchronizer flops themselves or constrain them. ATPG can then test reset logic by pulsing the test reset in capture.
Further reading: Siemens Tessent blog, hierarchical coverage reports
Core wrapping and hierarchical DFT
Isolating a core with wrapper cells so it can be tested on its own, generating patterns once at core level, and reusing them at the top. This is how every large SoC gets through ATPG.
60-second answer
A wrapper puts a controllable/observable cell on every core I/O. In internal mode the input cells drive the core and the output cells capture its response, so the core's patterns don't depend on the rest of the chip. That lets us run ATPG per core (fast, small memory), sign off coverage at core level, and retarget those patterns to the top through whatever channels the top gives the core. In external mode the wrapper cells test the top-level glue and wiring between cores. At the top, each wrapped core is replaced by a graybox model that only keeps the wrapper cells and the logic between them and the ports. Control is via IJTAG or an IEEE 1500 WIR.
Why wrap
- ATPG capacity: flat ATPG on a 108-gate SoC is slow and memory-hungry. Per-core runs parallelize and finish early in the schedule.
- Reuse: the same core instantiated 16 times gets one pattern set, broadcast to all instances.
- Schedule: core DFT signs off before the top is ready. Late top changes don't force re-ATPG of cores.
- Test scheduling and power: cores can be tested in groups to fit channel and power budgets.
- Diagnosis: failures localize to a core.
Wrapper cells
| Dedicated cell | Shared cell | |
|---|---|---|
| What | New flop + muxes on the port | Existing functional boundary flop doubles as wrapper cell |
| Area | Higher | Low |
| Timing | Adds mux on functional I/O path | No extra mux on the port path, but the flop's capture behavior is constrained in each mode |
| When | Unregistered ports, timing-relaxed I/O | Registered I/O (most well-designed cores); tools analyze fan-in/fan-out cones to find candidates |
Ports that don't need a wrapper cell: clocks, resets handled separately, test control pins, and ports whose logic is covered another way. Wrapper cells for a port feeding many flops through logic save a lot versus making every flop a shared cell. Tools apply a threshold on fan-in/fan-out cone size to choose.
Internal and external modes
| Mode | Input wrapper cells | Output wrapper cells | Tests |
|---|---|---|---|
| Internal (INTEST) | Drive core inputs (launch). Isolate from outside. | Capture core outputs (observe); drive safe values outward | Core logic |
| External (EXTEST) | Capture from outside (observe) | Drive outward (launch) | Glue logic, inter-core wiring |
| Functional | Transparent | — | |
The hierarchical flow
- Core level: insert wrapper cells, core EDT, OCC if needed, IJTAG-controlled mode signals. Run internal-mode ATPG, sign off coverage. Write a graybox (interface) model.
- Top level, internal: retarget core patterns to top-level pins. The tool maps core scan ports through top-level connections (channel sharing, broadcast, pipelines) and merges patterns of cores tested together.
- Top level, external: load grayboxes of all cores plus top logic; run ATPG with every core in external mode. Tests top glue and wiring.
- Coverage roll-up: sum core internal + top external fault lists to report the chip number (Tessent can merge flat fault lists across the hierarchy).
analyze_graybox/write_design -graybox-style commands; pattern retargeting maps core-level patterns to chip pins. Synopsys calls wrapped core models CTL models and does the same with DFT Compiler / TestMAX.IEEE 1500 in one picture
Summary only; the full treatment (serial control timing, every instruction, TAP/IJTAG control, CTL) is on the IEEE 1500 page.
- WSP
- Wrapper Serial Port: WSI, WSO plus the WSC control signals WRCK, WRSTN, SelectWIR, ShiftWR, CaptureWR, UpdateWR (TransferDR is optional in the standard).
- WIR
- Wrapper instruction register: selects mode and which register sits between WSI and WSO.
- WBY
- 1-bit bypass register, so you can shift through many wrapped cores quickly.
- WBR
- Wrapper boundary register: the chain of wrapper cells.
- Instructions
- WS_BYPASS and WS_EXTEST are mandatory, plus at least one internal-test instruction. Others include WS_INTEST_SCAN, WS_INTEST_RING, WS_PRELOAD, WS_CLAMP, WS_SAFE, and parallel-port (WP_) variants.
- vs 1149.1
- 1500 has no TAP state machine: the control signals are the decoded states (shift/capture/update). It's usually driven from a TAP or IJTAG network. Many teams now skip a full 1500 WIR and control wrappers directly with IJTAG TDR bits. Describing it as "1500-style wrapper, IJTAG-controlled" is accurate for most modern flows.
Interview questions
What's a graybox model and why is it enough for top-level ATPG? core
A reduced netlist of a wrapped core: wrapper chains, the logic between ports and wrapper cells, and the control/clock logic needed to operate them. The internal logic is removed. External-mode ATPG only launches from and captures into wrapper cells, so that's all it needs. It shrinks the top-level design by orders of magnitude.
How do identical core instances share test resources? core
Channel broadcast: the same compressed input stream goes to all instances, and each instance's outputs are compared (or compacted) separately, so any instance can fail independently. Pattern count is that of one core. Care: instances must be in the same mode and clocked identically, and output pins must be enough to observe each (or use on-chip compare).
A top-level port of a core is unregistered and feeds 2,000 flops through logic. Shared or dedicated cell? core
Dedicated. With no single boundary flop, a shared-cell approach would need to make every flop in the fan-out cone a wrapper cell (2,000 cells worth of constraints). One dedicated cell on the port isolates the whole cone.
Why can core-level coverage and chip-level coverage differ? deep
Retargeting constraints: at the top, the core may get fewer channels, pins tied, or clocks controlled differently (different OCC setup), so some patterns can't be applied as generated. Also faults on the core boundary are only covered in external mode, and top glue coverage depends on graybox accuracy. Report both and reconcile the difference.
How are wrapper modes controlled in your flow? basic
Either a 1500 WIR (instruction decode drives int/ext enables) or, more commonly now, IJTAG TDR bits in the core's ICL network. The test setup (PDL) sets those bits before scan; ATPG treats them as constant during the pattern set.
Further reading: Tsiatouhas, IEEE 1500 SECT lecture notes · Let's talk about wrappers (IEEE 1500)
Test point insertion
Adding a small amount of logic to make hard-to-control or hard-to-observe nets easy. It's essential for LBIST and increasingly used to cut deterministic pattern counts.
60-second answer
Some logic is random-pattern resistant: a wide AND needs all inputs at 1, which random patterns almost never produce, and a net deep in reconvergent logic may rarely propagate to a flop. Test points fix that. A control point adds an AND/OR gate driven by a test-point scan cell to force a net to 0 or 1 in test mode. An observe point taps a net into a scan cell. Tools pick locations with testability analysis (SCOAP-style or probabilistic) and avoid timing-critical paths. For LBIST, test points are what get you from ~80% to 90%+ coverage. For deterministic ATPG with compression, test points aimed at pattern-count reduction can cut patterns by 2–4×.
Random-pattern resistance
Test point types
- Control-0 / control-1: force a net low/high. The TP scan cell supplies the value (random in LBIST, chosen by ATPG in deterministic mode). Gated by a global test-point enable so functional mode is unaffected.
- Inversion (XOR) control point: flips a net instead of forcing it. Less biasing of the signal distribution.
- Observe point: net → dedicated scan cell or XOR tree into a scan cell. No gate in the functional path, only a load on the net. Cheapest in timing.
- Sharing: one TP scan cell can drive several control points, and XOR trees compress many observe points into one cell.
Testability metrics (SCOAP)
SCOAP gives each net a combinational 0-controllability (CC0), 1-controllability (CC1) and observability (CO), counting roughly how many assignments are needed. Higher is harder. PIs and scan cell outputs start at 1; POs and scan cell inputs have CO = 0.
OR: CC0(z) = Σ CC0(inputs) + 1 CC1(z) = min CC1(inputs) + 1
NOT: CC0(z) = CC1(a) + 1, CC1(z) = CC0(a) + 1
CO(input a of AND) = CO(z) + Σ CC1(other inputs) + 1 (side inputs at non-controlling value 1)
For the 16-input AND built from 4-input ANDs, CC1 at Z = 4×(4×1+1)+1 = 21, while CC0 = 1+1+1 = 3. That asymmetry is what flags it for a control-1 point. Commercial tools use probabilistic (COP-like) estimates plus fault simulation to rank candidate locations by expected coverage gain.
Where test points are used, and what they cost
| Use | Goal | Notes |
|---|---|---|
| Logic BIST | Raise random-pattern coverage | Essential. LBIST relies on PRPG + MISR, no deterministic top-off at runtime, so RP-resistant faults need TPs. Used for in-system/automotive test (ISO 26262). |
| EDT / deterministic | Cut pattern count | TPs that make faults easier to detect together reduce patterns (Tessent offers pattern-count-targeted TPI). Also helps when compression encoding is tight. |
| Hybrid LBIST+EDT | Share logic | Same TPs serve both; common in automotive chips. |
| X-bounding | Block X sources | Control points on X sources (memories, black boxes) so they don't corrupt MISR/compactor. |
- Timing: control points add a gate on a functional path. Tools exclude critical paths (read SDC/timing slack) and prefer observe points.
- Area: each TP costs a gate + share of a scan cell. Typically 1–2% of flops.
- Power: control points change toggle activity in test; enable only in test mode.
- Verification: equivalence check with TP enable = 0 must pass against the pre-TP netlist.
Interview questions
Why do control points hurt timing but observe points don't? basic
A control point inserts a gate in series on the net, adding delay to every functional path through it. An observe point only adds a fanout branch (capacitive load) to a scan cell; the functional path is unchanged apart from the extra load.
Why does LBIST need test points and ATPG mostly doesn't? core
ATPG targets each fault deterministically: it will set all 16 inputs of an AND to 1 if needed. LBIST applies pseudo-random patterns, so the probability of hitting such a combination is tiny. Test points change the circuit so random patterns detect those faults with reasonable probability.
How can TPI reduce deterministic pattern count? deep
Patterns are often limited by conflicts: two faults need incompatible values on some net, or a fault needs so many care bits that few other faults fit in the same pattern. Control points break those conflicts and observe points shorten propagation paths (fewer care bits per fault). More faults fit per pattern, so fewer patterns are needed. Useful when tester memory or test time is the constraint.
Where would you not put a test point? core
On timing-critical paths, clock and reset nets, asynchronous crossings, analog interfaces, and anything in a safety mechanism's functional path where an added gate would need extra safety analysis. Also not in paths where a forced value in test could cause contention or damage.
Low-power scan
Scan toggles far more logic than functional operation. Shift power heats the die and drops the supply; capture power causes IR-drop that slows paths and fails good parts at speed. How to control both.
60-second answer
There are two problems. Shift power is average power over long shift sequences: every flop toggles with random data, so switching activity can be several times functional. The fixes are slower shift, low-transition fill, Q-gating, and shifting only some chains or cores at a time. Capture power is the instantaneous peak at launch/capture in at-speed tests. IR-drop there slows paths, so good dies fail (yield loss) or bad dies escape if you relax the clock. The fixes are clock-gating-aware and power-aware ATPG that caps toggles per pattern, staggering domain clocks, and limiting how many cores capture together. With UPF domains you also keep chains inside one domain, test the isolation and retention cells, and test with domains in their valid power states.
Why test power is different
| Shift | Capture (at-speed) | |
|---|---|---|
| Nature | Average, sustained (thousands of cycles) | Peak, 1–2 fast cycles |
| Cause | Random data rippling through every flop + the logic behind them | Launch transitions across many domains simultaneously |
| Effect | Thermal, package/probe current limits, supply droop → shift failures | Dynamic IR-drop → slower paths → false at-speed fails, or reduced test frequency |
| Typical knobs | fshift, fill, Q-gating, chain/core partitioning | Power-aware ATPG, clock gating usage, staggered/sequential capture, fewer active domains |
Shift power
- Lower shift frequency: power ∝ f. Simplest, but test time goes up proportionally.
- Clock gating in shift: shift only part of the design at a time (per core or per chain group), holding the other groups' clocks off. Costs test time.
- Q-gating: blocks the combinational logic from seeing shift toggles. Very effective on logic power, costs area and functional timing.
- Low-power EDT: the decompressor can hold values for chains that have no care bits in a given shift window, which mimics adjacent fill.
- Scan chain ordering: grouping flops so adjacent cells tend to hold equal values reduces toggles a bit. Minor compared to fill.
Capture power
scan_en and ATPG controls the functional enable during capture (tests the enable logic, lower capture power); tie it to test_mode and the clock is always on in test (the enable logic becomes untestable and capture power goes up).- Power-aware ATPG: set a budget (e.g. maximum % of flops that change state at capture, or weighted switching activity). ATPG respects it using fill choices and by turning off clock gates it doesn't need (enables = 0) for that pattern. More patterns result.
- Clock-gating aware: requires TE tied to SE, not test_mode, so ATPG controls the functional enables during capture (see figure).
- Staggered capture: domains pulse at different times in the capture window rather than together. The OCC per domain supports this.
- Core scheduling: in hierarchical test, choose which cores capture together to meet the peak budget.
- Validation: run vector-based IR-drop (e.g. RedHawk/Voltus with the actual patterns' switching) on worst patterns before tapeout; correlate with silicon.
Power domains and scan
- Chains don't cross switchable domains. If a domain is off, any chain passing through it breaks. Build chains per domain (tools read UPF during insertion).
- Isolation cells clamp outputs of an off domain. Test their clamp behavior with the domain in the off state, and that they're transparent when on.
- Level shifters between voltage domains may sit on scan paths; they add delay and need timing in shift mode across voltages.
- Retention flops: need save/restore testing, often a dedicated retention test sequence (save, power down, power up, restore, unload).
- Power states in test: patterns run with a defined power-state table (which domains on, which voltages). Test all valid states that matter, not just all-on.
- Power switch control must be controllable in test mode and not toggle during shift.
Interview questions
Why can test power exceed functional power? basic
Functional workloads have correlated data and most logic idle; typical toggle rates are low. Scan data is essentially random and every flop shifts every cycle, and at capture ATPG deliberately activates many paths at once. Also functional designs rely on clock gating to keep idle blocks off, which test may bypass.
Good dies fail transition patterns at nominal voltage but pass at +5% V. What do you suspect? core
Capture IR-drop. Too many transitions at launch drops local supply, slowing paths. Extra voltage restores speed. Confirm: check the failing patterns' switching activity; rerun with a power-aware pattern set (toggle cap) and see if fails go away; compare fails vs path slack. Fix with power-aware ATPG, staggered capture, or fewer simultaneous domains.
How does adjacent fill interact with compression? core
With a linear decompressor you don't choose non-care bits freely; they're whatever the decompressor produces, which looks random. Low-power compression modes add hold logic so that chains (or groups) with no care bits in a window receive a constant, which recreates the adjacent-fill effect at a cost of some encoding capacity.
What's wrong with tying ICG TE to test_mode? core
The clock is always on in test, so (1) the functional enable logic can't be tested (its effect is masked); (2) every gated flop captures every pattern, raising capture power; (3) you lose the ability to use clock gating for power-aware ATPG. With TE = SE the gate is forced on only during shift, and ATPG controls the enable in capture.
JTAG (IEEE 1149.1)
The Test Access Port: five pins, a 16-state controller, an instruction register and a set of data registers. It's the front door for boundary scan, IJTAG networks, MBIST, test-mode entry and silicon debug.
60-second answer
1149.1 defines a serial port (TCK, TMS, TDI, TDO, optional TRST*) and a 16-state TAP controller driven by TMS on rising TCK edges. There are two symmetric paths: load an instruction through the IR, then that instruction selects which data register sits between TDI and TDO. Each scan goes Capture → Shift → Update, so the register's parallel outputs only change at Update and don't ripple while shifting. Mandatory pieces are the IR, BYPASS, the boundary-scan register and the instructions BYPASS, SAMPLE, PRELOAD and EXTEST. On an SoC it's mostly used as the access mechanism for internal test infrastructure through user-defined data registers and IJTAG.
The TAP controller
- Reset
- TMS=1 for 5 TCKs reaches Test-Logic-Reset from any state (or assert TRST*). In TLR the IR holds IDCODE if the device has an ID register, else BYPASS, and test logic is inactive.
- Capture-xR
- Parallel load into the shift stage on the rising edge leaving the state. IR capture loads
…01into the two LSBs (the other bits are design-specific), which lets software check the chain. - Shift-xR
- One bit per rising edge, LSB first out of TDO. The edge that moves from Shift to Exit1 also shifts, so an n-bit register needs TMS=0 for n−1 edges in Shift and TMS=1 on the last one.
- Pause-xR
- Hold shifting without losing state, e.g. while the tester reloads buffers.
- Update-xR
- Shift stage copied to the parallel hold stage on the falling edge of TCK.
- Run-Test/Idle
- Idle state, and where self-tests run (RUNBIST, MBIST started via a TDR and left running for N TCKs).
Architecture and instructions
| Instruction | Status | Register between TDI/TDO | Effect |
|---|---|---|---|
| BYPASS | Mandatory, all-1s opcode | BYPASS (1 bit, captures 0) | Shortest path through a chip on a board chain |
| SAMPLE | Mandatory | BSR | Capture pin values without disturbing function |
| PRELOAD | Mandatory (may share opcode with SAMPLE) | BSR | Load BSR update stage before EXTEST so pins start in a known state |
| EXTEST | Mandatory | BSR | Pins driven from BSR; board interconnect test. Historically all-0s opcode, no longer required since 2001 |
| IDCODE | Optional (required if ID register exists) | 32-bit ID | LSB=1, 11-bit JEDEC manufacturer, 16-bit part, 4-bit version |
| INTEST, RUNBIST, CLAMP, HIGHZ | Optional | BSR / BYPASS | Core test via BSR; self-test; hold pins while bypassed; tri-state all outputs |
| USER / private | Design-defined | Any TDR | IJTAG host, MBIST, fuse, PLL, test-mode control |
Timing
Boundary-scan cell
The two-stage structure (capture/shift + update) is the idea to remember. 1500 wrapper cells and 1687 TDRs reuse it: shifting never disturbs parallel outputs, and everything updates at once.
Interview questions
From Run-Test/Idle, what TMS sequence shifts 8 bits into a DR and returns to Idle? core
RTI → Select-DR (1) → Capture-DR (0) → Shift-DR (0). Then 7 more rising edges with TMS=0 in Shift (7 bits), and the 8th bit shifts on the edge with TMS=1 into Exit1-DR. Then Update-DR (1), then RTI (0). TMS: 1 0 0 | 0000000 1 | 1 0 = 13 TCKs.
Why does IR capture load 01 in the LSBs? core
Chain integrity. When scanning a board's IR chain, each device shifts out its fixed 01 pattern first, so software can check the chain is connected and count devices and IR lengths. The 1 guarantees a 0→1 transition, which detects stuck-at-0 and stuck-at-1 on TDO.
Why does TDO change on the falling edge? basic
So the next device in a daisy chain (or the tester) samples on the rising edge with half a TCK period of setup and hold margin. It makes chaining devices across a board robust to skew.
How do you find the number of devices on an unknown JTAG chain? core
Reset (all IRs → IDCODE or BYPASS). Put everything in BYPASS by shifting a long string of 1s into the IR chain. Then flush the DR chain with 0s and shift a 1: the number of TCKs until the 1 appears at TDO equals the number of 1-bit BYPASS registers, i.e. the device count. Alternatively after reset, read the DR chain: IDCODE registers start with 1, bypass captures 0.
SAMPLE vs PRELOAD vs EXTEST? basic
SAMPLE captures live pin values into the BSR while the chip runs normally. PRELOAD shifts a known pattern into the BSR update stage, still in normal mode. EXTEST then hands pin control to the BSR, driving those preloaded values out and capturing what arrives, which tests board interconnect. PRELOAD first avoids driving garbage when EXTEST activates.
What does 1149.1 give an SoC DFT engineer beyond boundary scan? core
A standard, always-available, low-pin-count access port. Through private instructions and TDRs you enter test modes, program OCC/PLL settings, start and read MBIST, blow/read fuses, configure scan and compression modes, and in silicon debug read status registers. Today mostly as the host for an IJTAG network.
Further reading: Corelis JTAG technical primer · Electronic Design, JTAG vs IJTAG
IJTAG (IEEE 1687)
A standard way to describe and access hundreds of on-chip instruments through a reconfigurable scan network, and to write their test procedures once at the instrument level and have a tool retarget them to the chip's pins.
60-second answer
1149.1 gives you a port and a few data registers, but an SoC has hundreds of instruments: MBIST controllers, PLL and OCC configuration, sensors, wrapper enables, fuse controllers. 1687 defines a scan network behind the TAP built from TDRs, SIBs and scan muxes, so the active path contains only what you're accessing. It has two languages. ICL describes the network's structure. PDL describes operations on an instrument, like "write this field, run 1,000 clocks, read that status bit". A tool combines the two to compute the TAP-level vectors (retargeting). Instruments and their PDL are reusable across designs and hierarchy levels, and the tool handles path lengths and SIB sequencing.
Why 1687 exists
- Scale: one flat user DR per instrument doesn't scale; one long chain containing everything is slow and fragile.
- Reuse: an instrument's access procedure (PDL) is written against its own ports and registers, not against a specific chip.
- Hierarchy: a core's network plugs into the parent's network; procedures stay valid.
- Automation: the tool, not a person, figures out SIB states, path lengths and the order of scans.
The network
- TDR
- Test data register: a shift stage and (usually) an update stage, like a 1149.1 DR. Its parallel outputs drive instrument inputs; its capture inputs read instrument status.
- SIB
- Segment insertion bit: 1-bit register that inserts or bypasses a hosted segment. Builds hierarchical, variable-length networks.
- ScanMux
- Selects between scan segments, controlled by bits elsewhere in the network (e.g. a TDR field). More general than a SIB.
- AccessLink
- ICL construct connecting the network to the 1149.1 TAP (which instruction opens it). The network can in principle be reached through other ports too.
- Instrument
- The thing being controlled (MBIST, sensor, PLL). It sees only parallel data ports, never TCK/TMS.
The SIB in detail
ICL and PDL
ICL describes structure: modules, scan in/out ports, scan registers, muxes, data ports and how they connect. PDL describes operation at the instrument's level.
// ICL (simplified): a TDR driving a PLL's controls and capturing its lock status
Module pll_ctrl_tdr {
ScanInPort si; ScanOutPort so { Source ST[0]; }
ShiftEnPort se; CaptureEnPort ce; UpdateEnPort ue;
SelectPort sel; TCKPort tck;
DataOutPort div_o[3:0] { Source R[7:4]; }
DataOutPort byp_o { Source R[1]; }
DataInPort lock;
ScanRegister R[7:0] { ScanInSource si; ResetValue 8'h00; }
ScanRegister ST[0:0] { ScanInSource R[0]; CaptureSource lock; }
Alias div[3:0] = R[7:4];
Alias bypass = R[1];
}
# PDL (Tcl-based, Level-1): configure the PLL and check lock
iProcsForModule pll_ctrl_tdr
iProc set_div {ratio} {
iWrite div $ratio ;# queue writes
iWrite bypass 0
iApply ;# one retargeted scan (or more)
iRunLoop 2000 -tck ;# wait for lock
iRead ST 1 ;# expect captured lock = 1
iApply
}
- iWrite / iRead
- Queue a write, or an expected read (compare value), on a register field or data port.
- iApply
- Execute everything queued. The tool may merge operations on different instruments into one scan.
- iScan
- Raw scan of a register with given data (less common).
- iRunLoop
- Wait N clock cycles (TCK or a named system clock), e.g. for MBIST to run.
- iProc / iCall
- Define and invoke reusable procedures, scoped to a module. Calls are hierarchical:
iCall coreA.pll.set_div 4. - iClock
- Declare that a system clock must be running for the following operations.
- iMerge
- Allow operations from parallel threads to share scans.
Retargeting
Given ICL for the whole chip and PDL at instrument level, the tool:
- Finds the path from the TAP to each target register in the ICL graph.
- Computes the SIB/ScanMux settings needed and the sequence of scans to reach them (opening outer SIBs first).
- Assembles each scan's bit vector with the right length, placing writes and expected reads at the right offsets. Other bits keep their values.
- Emits TAP-level patterns (TMS/TDI/expected TDO) for the tester or a testbench, or drives them through a DFT test-setup procedure before scan ATPG.
iCall in a test-setup procedure), so the scan patterns inherit a known DFT configuration. When silicon doesn't enter a mode, simulate the retargeted setup vectors first.Interview questions
What's the difference between ICL and BSDL? basic
BSDL describes a chip's 1149.1 boundary-scan implementation for board test (pins, BSR cells, instructions). ICL describes the internal instrument network behind the TAP: registers, SIBs, muxes, and how instruments connect. They complement each other.
Why do SIBs make access faster? core
With everything closed the active path is a handful of bits. You open only the branch you need, so each scan is short. A flat daisy chain of every TDR would make every access as long as the sum of all registers, often tens of thousands of bits.
A retargeted PDL sequence works in simulation but silicon reads back all zeros. Where do you look? deep
First TAP basics: IDCODE reads correctly? Does the IJTAG host instruction load? Then network integrity: shift a known pattern through the closed network (just SIB bits) and check the length. Then open SIBs one level at a time and check each length. All zeros suggests a broken path (so TDO stuck), a wrong instruction opcode, or TCK/reset issues. Check TRST/reset values of SIBs and that the tester's TDO strobe timing is correct (falling-edge launch).
How does IJTAG relate to 1500? core
1500 standardizes the core wrapper and its serial control. 1687 standardizes the access network and procedure description. In practice wrapper controls are often just TDR bits in the core's 1687 network, and 1500-style wrappers are described in ICL. They coexist: 1687 can host a 1500 WIR as a register.
What does iApply guarantee about ordering? core
All operations queued since the last iApply complete before the next one starts. Within one iApply, the tool may combine writes to different registers into one scan and may need several scans if SIB reconfiguration is required. Reads compare what's captured in that scan. If you need a write, then time, then a read, you need separate iApply calls with an iRunLoop between them.
Further reading: ASSET InterTech, ICL and PDL explained · Electronic Design, scan ATPG vs IJTAG retargeting · Nordic Test Forum, IEEE 1687 slides
IEEE 1500 core wrapper
The standard for isolating an embedded core so it can be tested on its own: a boundary register around the core, an instruction register that sets the wrapper's mode, and a serial port that looks like a TAP without the state machine.
60-second answer
IEEE 1500 (Standard Testability Method for Embedded Core-based ICs) standardizes the wrapper around a reusable core. It has four register pieces: the WBR (a cell on every functional I/O), the WIR (instruction register), the WBY (1-bit bypass), and the core's own data registers such as scan chains. They're reached through the Wrapper Serial Port: WSI and WSO plus six control signals (WRCK, WRSTN, SelectWIR, ShiftWR, CaptureWR, UpdateWR). An optional parallel port (WPI/WPO) carries high-bandwidth scan. Unlike 1149.1 there's no state machine on the core. The control signals are already decoded, so one chip-level TAP (or IJTAG network) can drive many wrappers. Mandatory instructions are WS_BYPASS and WS_EXTEST plus at least one internal-test instruction. The core's test interface is described in CTL (IEEE 1450.6), so the integrator can reuse the core's patterns.
What problem 1500 solves
- Isolation: the core can be tested without knowing what surrounds it, and the surroundings (glue logic, other cores) can be tested without the core.
- Reuse: the core provider ships the wrapper plus patterns and a CTL description; the SoC integrator retargets the patterns to chip pins.
- Plug-and-play access: a standard serial control interface, so a chip-level controller can manage dozens of wrappers the same way.
- Scalability: serial access for control and low-bandwidth test, parallel ports for real scan bandwidth.
1500 standardizes the wrapper and its interface. It deliberately does not standardize the chip-level access mechanism (TAM): how WPI/WPO and the WSP reach the pins is the integrator's choice.
Architecture
- WSP
- Wrapper Serial Port: WSI, WSO, and the WSC control signals WRCK, WRSTN, SelectWIR, ShiftWR, CaptureWR, UpdateWR. TransferDR is an optional extra control, and the standard also allows auxiliary clocks (AUXCK).
- WPP
- Wrapper Parallel Port (optional): WPI, WPO and WPC (parallel control). This is how real scan data reaches the core: many channels into core chains or EDT.
- WIR
- Instruction register, with a shift stage and an update stage. Its decoded value sets every WBR cell's mode and picks the register between WSI and WSO. WRSTN resets it asynchronously so the wrapper is transparent (functional mode).
- WBY
- 1-bit bypass, so a chain of many wrapped cores stays short when you only want one.
- WBR
- Boundary register made of wrapper boundary cells, one per functional port bit (clocks, resets and test pins are typically excluded).
- Core data registers
- Scan chains, MBIST or other core registers that instructions can put on the serial or parallel path.
Serial control timing
Instructions
Naming: WS_ instructions use the serial port for test data, WP_ the parallel port. The standard mandates WS_BYPASS and WS_EXTEST plus at least one core-test (INTEST-type) instruction. The rest are optional or user-defined.
| Instruction | Status | WSI→WSO register | Wrapper behavior | Used for |
|---|---|---|---|---|
| WS_BYPASS | Mandatory | WBY | Cells transparent; core runs functionally. Default after WRSTN. | Normal mode, short path through untested cores |
| WS_EXTEST | Mandatory | WBR | Output cells drive outward from the WBR; input cells capture from outside. Core isolated. | Glue logic and interconnect between cores |
| Wx_INTEST (one required) | Mandatory (at least one) | depends | Input cells drive the core; output cells capture from it | Testing the core itself |
| WS_INTEST_SCAN | Optional form | WBR + core chains, concatenated | All serial through WSI/WSO | Low-pin-count or debug scan of the core |
| WS_INTEST_RING | Optional form | WBR only | Core stimulated and observed through its boundary only | Cores without internal scan (or functional/BIST-style tests) |
| WP_INTEST | Optional form | Parallel: WPI/WPO | Core chains and WBR segments on the parallel port | Production scan: the one that carries real bandwidth |
| WP_EXTEST | Optional | Parallel | EXTEST via the parallel port | Faster external test |
| WS_PRELOAD / WP_PRELOAD | Optional | WBR | Load cell update stages while the wrapper stays functional | Preset outputs before switching to EXTEST |
| WS_CLAMP | Optional | WBY | Outputs held at WBR values while the short WBY path is selected | Keep a core's outputs fixed while testing neighbors |
| WS_SAFE | Optional | WBY | Cells force predefined safe values on outputs | Isolation: stop an untested or powered-down core from disturbing others |
Wrapper boundary cells
- Events a cell supports: shift (along the WBR), capture (from its functional input), update (copy the shift stage to a hold stage, if the cell has one), and apply (drive its functional output from the stored value).
- Dedicated vs shared: a dedicated cell is a new flop added on the port. A shared cell reuses an existing functional boundary flop. Shared saves area and a mux on the path, but its capture behavior is constrained in each mode. See core wrapping.
- With or without update stage: an update stage keeps outputs steady while the WBR shifts (needed if downstream logic must not see rippling values, e.g. when driving other cores' inputs). Without it, the cell is smaller but outputs ripple during shift.
- Safe values: cells used by WS_SAFE need a way to force a fixed output value independent of the shift data.
- Naming: the standard defines a naming scheme for cell types (names like
WC_SD1_CII) that encodes the storage elements and where capture and update happen. It's worth recognizing; you won't be asked to decode it.
Driving wrappers from a TAP or an IJTAG network
- From a TAP: the TAP state machine produces ShiftWR/CaptureWR/UpdateWR, TCK becomes WRCK, and a TAP instruction asserts SelectWIR. Wrappers are daisy-chained; untested cores sit in WS_BYPASS.
- From IJTAG: each wrapper's WIR (or just its mode bits) is a TDR behind a SIB. Only the wrappers you open are in the path. This is how most modern flows do it. Tessent's hierarchical flow, for example, uses 1500-style wrapper cells whose internal/external mode enables are TDR bits in the core's IJTAG network.
- Parallel port at chip level: WPI/WPO connect to chip scan pins or to a shared test bus, often through muxes or channel broadcast, so several cores can be scanned at once.
- CTL (IEEE 1450.6)
- Core Test Language, an extension of STIL. It describes the core's test modes, wrapper structure, scan chains and pattern references, so the integrator's tools can connect and retarget without reading the core's netlist.
- Compliance
- The standard distinguishes a wrapped compliant core (wrapper included, described in CTL) from an unwrapped compliant core (no wrapper yet, but with CTL so a compliant wrapper can be added later). You'll also hear "1500-ready" used informally for the second.
1149.1 vs 1500 vs 1687
| 1149.1 (JTAG) | 1500 (core wrapper) | 1687 (IJTAG) | |
|---|---|---|---|
| Scope | Chip pins, board test, chip access port | Boundary of an embedded core | Network of on-chip instruments |
| Control | TMS + 16-state FSM | Decoded signals (WSC), no FSM | Uses a host port (usually the TAP); SIB/ScanMux configure the path |
| Instruction register | IR | WIR | None of its own; configuration bits live in the network |
| Boundary register | BSR on pins | WBR on core ports | — |
| Bypass | BYPASS (1 bit) | WBY (1 bit) | Closed SIB (1 bit) |
| Description language | BSDL | CTL (1450.6) | ICL + PDL |
| High-bandwidth path | No | Yes: parallel port (WPI/WPO) | No (access network; scan data goes elsewhere) |
| Typical use today | Chip access, boundary scan, IDCODE | Wrapper architecture concepts, core isolation, hierarchical ATPG | Accessing MBIST, PLLs, sensors, wrapper/EDT/OCC setup |
Interview questions
Why doesn't a 1500 wrapper have a state machine like the TAP? core
Because cores are embedded, the chip already has a controller (TAP or IJTAG) that decodes the protocol. Giving the wrapper decoded control signals (ShiftWR, CaptureWR, UpdateWR, SelectWIR) means one controller can drive many wrappers in parallel over a broadcast bus, with less area per core and no risk of per-core FSMs getting out of sync.
Which instructions are mandatory? basic
WS_BYPASS, WS_EXTEST and at least one internal-test instruction (a Wx_INTEST form, such as WS_INTEST_SCAN, WS_INTEST_RING or WP_INTEST). Everything else (WS_SAFE, WS_CLAMP, PRELOAD, parallel variants, user instructions) is optional.
WS_SAFE vs WS_CLAMP? core
Both select the 1-bit WBY so the serial path stays short. WS_CLAMP holds the outputs at whatever was loaded into the WBR (you choose the values, like 1149.1 CLAMP). WS_SAFE drives predefined safe values built into the cells, for example to isolate a core that's powered down or untested so it can't disturb its neighbors.
What's WS_INTEST_RING for, if WP_INTEST exists? deep
It tests the core only through its boundary: input cells apply values, output cells capture. It's useful for cores without internal scan (small or legacy cores, or hard macros tested functionally), or for applying a BIST-style or functional pattern. WP_INTEST is the high-bandwidth scan path for cores with internal chains.
How would you connect 20 wrapped cores to one TAP? core
Broadcast the WSC from the TAP (WRCK = TCK, Shift/Capture/UpdateWR from the TAP states, SelectWIR from a TAP instruction, WRSTN from reset) and daisy-chain WSI/WSO between TDI and TDO. First load all WIRs in one long instruction scan: the core under test gets its INTEST/EXTEST instruction and the rest get WS_BYPASS (or WS_SAFE). Then the data scan passes through one WBR plus 19 bypass bits. For better scaling, put each wrapper behind a SIB in an IJTAG network. For scan data, route WPI/WPO to chip scan pins, with muxing or broadcast.
What does WRSTN do, and why does it matter at power-up? basic
It's the asynchronous active-low wrapper reset. It forces the WIR to its functional state (bypass, cells transparent). If the WIR powered up in a random test mode, a core's outputs could be driven from garbage WBR values in functional operation. So WRSTN must be asserted at reset and tied off correctly in the chip.
Why do 1500 and IJTAG coexist in the same chip? core
They solve different problems. 1500 defines the wrapper around a core (boundary cells, modes, isolation). 1687 defines how you reach and configure on-chip instruments. In practice the wrapper's mode control is often just TDR bits in the core's IJTAG network, and the scan data still moves through parallel scan/EDT channels. You can describe a 1500 wrapper in ICL and drive its WIR as an IJTAG register.
What's in the CTL for a core, and who uses it? deep
Test modes and how to enter them, wrapper and scan structure (which ports are scan in/out, chain lengths, clocks), signal constraints per mode, and references to the core's test patterns. The SoC integrator's DFT tools use it to connect the core's test ports, check compatibility and retarget the core-level patterns to chip pins without the core's full netlist.
Further reading: Tsiatouhas, IEEE 1500 SECT lecture notes · Arm, IEEE 1500 compliant wrapper boundary register cell · Marinissen et al., Overview of the IEEE P1500 standard · Da Silva, McLaurin, Waayers, The Core Test Wrapper Handbook
Fault models and ATPG
What defects we model, how ATPG finds a pattern for each fault, how faults are classified, and how coverage is computed and debugged.
60-second answer
A fault model is an abstraction of physical defects that ATPG can target. Stuck-at covers static shorts and opens; transition faults cover gross delay defects with a launch-capture pair at speed; path delay and timing-aware ATPG go after small delays on long paths; bridging and cell-aware models target specific layout or intra-cell defects. For each fault, ATPG excites it (sets the site opposite the stuck value), propagates the difference to an observable scan cell or PO through non-controlling side inputs, and justifies all required values back to scan cells and PIs. Fault simulation then drops every other fault the pattern detects, and dynamic compaction packs more faults into each pattern. Test coverage is detected over testable faults, with untestable faults excluded. The ATPG-untestable faults are the ones you debug: they point to constraints, black boxes, clocking and X.
Fault models
| Model | Targets | Test | Notes |
|---|---|---|---|
| Stuck-at (SA0/SA1) | Static shorts to rail, opens | One pattern, slow capture | 2 faults per pin; baseline, ≥99% TC typical target |
| Transition (STR/STF) | Gross delay at a node | Two patterns: V1 init, V2 launch; capture one period later | LOC/LOS; ~90–95%+ TC typical; lower than SA because of clocking/exception limits |
| Path delay | Cumulative delay on a specific path | Two patterns sensitizing the whole path | Robust vs non-robust sensitization; used on critical paths from STA |
| Small delay defect (timing-aware) | Small extra delays | Transition patterns through longest paths (low slack) | Uses SDF/slack; more patterns |
| Bridging | Shorts between neighboring nets | Drive nets to opposite values, observe victim | Net pairs extracted from layout |
| Cell-aware | Defects inside standard cells | Cell-specific input combinations (static and delay) | From transistor-level defect simulation of each library cell; catches what SA misses |
| IDDQ | Leakage-causing defects (bridges) | Measure quiescent current | Less effective at advanced nodes (high background leakage) |
Fault collapsing
An n-input AND has 2n+2 stuck-at faults on its pins. Any input SA0 is equivalent to output SA0 (same tests detect them), so they collapse to n+2 classes. Dominance: output SA1 of an AND is detected by any test for an input SA1, so it can be dropped when targeting. Checkpoint theorem: tests for stuck-at faults on PIs and fanout branches detect all single stuck-at faults in a fanout-free structure. Tools report both collapsed and uncollapsed counts; know which one a coverage number uses.
ATPG algorithms
- D-algorithm (Roth)
- 5-valued algebra (0, 1, X, D, D̄). Propagates the D-frontier forward, justifies the J-frontier backward; decisions on internal nodes.
- PODEM (Goel)
- Decides only on primary inputs (and scan cells), backtraces objectives to an input, then simulates. Smaller search space and far fewer backtracks than D-alg.
- FAN
- PODEM plus decisions at headlines and fanout stems, multiple backtrace. Modern tools add learning (static and dynamic implications) and SAT solvers for hard faults.
- Backtrack limit
- Give up on a fault after N backtracks: it's aborted. Raise the limit or add effort in a second pass to resolve aborts.
- Fault simulation
- After each pattern, simulate all remaining faults (bit-parallel or concurrent) and drop the ones detected fortuitously. Most faults are detected this way, not by being targeted.
- Compaction
- Dynamic: after generating a test for one fault, fill remaining don't-cares to detect more targeted faults in the same pattern. Static: afterward, merge compatible patterns and drop redundant ones (reverse-order fault sim).
- N-detect
- Detect each fault N times through different patterns; improves unmodeled defect coverage at the cost of more patterns.
Transition faults
Why transition coverage is lower than stuck-at: a pattern needs a legal launch (LOC can only launch transitions the logic produces from V1), both launch and capture domains must be pulsed with a known relationship, paths through false/multicycle exceptions must be masked, and non-scan or X sources hit twice. Two-cycle sequential depth also makes ATPG harder. Typical numbers: SA ~99%, TDF ~90–97% depending on design and clocking.
Fault classes and coverage
fault coverage = (DT + c·PT) / total
c = possible-detect credit (commonly 50%)
ATPG effectiveness = (DT + UT + AU + PU + c·PT) / total (how many faults ATPG resolved)
Test coverage is what you sign off. It's the fraction of faults that can be tested that are tested. Fault coverage penalizes redundant/tied logic you can't do anything about. Always state which one, collapsed or not, and at which level of hierarchy.
Coverage debug: from AU to action
| Cause | Symptom | Fix |
|---|---|---|
| Pin constraints / test-mode ties | Logic controlled by a constrained pin is blocked | Relax constraints; test the logic in another mode |
| Black boxes / RAMs | Shadow logic AU (inputs unobserved, outputs X) | Bypass, observe points, RAM sequential ATPG, or model |
| Non-scan cells | X at capture, uncontrollable state | Scan them, initialize them, or add test points |
| Clock restrictions | Faults between domains not pulsed together | Allow multi-domain capture where synchronous; add capture procedures |
| Timing exceptions (at-speed) | Endpoints of false/MCP masked | Expected; verify exceptions are correct and not too broad |
| X sources + compression | Faults observable only through heavily masked chains | Remove X at source; more masking granularity; X-bounding TPs |
| Aborted (UD) | Hard faults | Higher backtrack limit, second pass with more effort |
report_statistics for the class breakdown, report_faults -class AU (and sub-classes) to list, analyze_fault on a sample to see why it's untestable, and grouping AU by instance to find the one block costing you 0.5%.Interview questions
Test coverage is 98.2%. How would you find the missing 1.8%? core
Break down the non-detected faults by class (AU, UD, PT/PU) and then by hierarchy/instance. Usually a few blocks dominate. For each: check pin constraints and test-mode ties, black boxes and memories, non-scan flops, clock-domain capture restrictions, and X masking. Run fault analysis on samples to get the actual blocking reason. Separate "fixable in ATPG setup" from "needs design change" (test points, bypass logic) from "expected" (redundant).
Why is PODEM more efficient than the D-algorithm? core
PODEM only makes decisions on primary inputs/scan cells, so every decision is followed by simulation that gives consistent values; there are no internal-node inconsistencies to discover late. The search space is 2#inputs rather than over all internal nodes, and backtracking is simpler.
What's a redundant fault? Give an example. core
A fault that no input combination can detect because the faulty and good circuits are logically equivalent. Example: z = a·b + ā·c + b·c. The b·c term is the consensus term: whenever b·c = 1, either a·b or ā·c is already 1. So the output of the b·c AND stuck-at-0 never changes z and can't be detected. Consensus terms like this are often added on purpose to avoid hazards. RE faults are excluded from test coverage; the logic could be removed.
Why do you need both stuck-at and transition patterns if transition patterns also detect stuck-at faults? core
Transition patterns do detect many stuck-at faults, and many flows fault-simulate transition patterns against stuck-at to top off. But stuck-at patterns are cheaper (single capture, fewer constraints), can use slow capture on logic that has no at-speed clocking, and reach faults blocked in at-speed mode (masked exception endpoints, domains not pulsed together). So: run transition first, fault-sim for SA credit, then top off with SA ATPG.
What is posdet (possibly detected)? basic
The good machine produces a known value at an observation point but the faulty machine produces X, so detection depends on what the X resolves to in silicon. Tools give partial credit (often 50%). Commonly caused by faults on clock or enable lines, or feeding X sources.
What is cell-aware test and why is it worth the extra patterns? deep
Each library cell is characterized at transistor level: inject opens/shorts in the layout, simulate, and record which input combinations (static and two-cycle) detect each defect. ATPG then targets those specific combinations. Stuck-at at cell pins doesn't guarantee these combinations are applied, so cell-aware catches intra-cell defects that escape; industry data shows measurable DPPM reduction. Cost: more patterns and library characterization effort.
Further reading: Fault classes, Mentor terminology · VLSI Space, test vs fault coverage · Wikipedia, fault coverage
Gate-level simulation and SDF
Simulating ATPG patterns on the real netlist, with and without back-annotated timing, to prove they'll pass on good silicon. Also the systematic way to debug a mismatch back to its root cause.
60-second answer
ATPG computes expected values with its own zero-delay model and its own understanding of the test setup. GLS checks that against a real simulator on the final netlist. Zero-delay GLS catches functional issues: wrong test-mode setup, IJTAG/OCC programming, clock muxing, X from uninitialized state, library model mismatches. SDF-annotated GLS adds the real delays and timing checks at each corner (max/setup at slow, min/hold at fast) and catches shift hold races, SE timing, capture timing and glitches that STA constraints may have missed or mis-modeled. Parallel-load sim runs all patterns fast; serial sim runs a subset to prove the chains. To debug a mismatch I go from the first failing pattern and cycle, map it to the scan cell, trace back in the waveform to the first X or wrong value, and then decide whether it's timing, setup or a model issue.
What GLS catches that ATPG and STA don't
- Test setup: IJTAG/TDR programming, test-mode entry sequence, PLL lock time, OCC configuration. ATPG often assumes these (via test procedures); simulation proves them.
- Constraint gaps: a path STA didn't time in the right mode (missing case analysis, over-broad false path), or a clock mux STA assumed static.
- X and initialization: non-reset flops, memories, analog models that go X in simulation.
- Library issues: Verilog model vs Liberty vs ATPG library disagreeing (e.g. scan cell pin function, ICG behavior).
- Glitches on clocks or async resets from muxing or gating in test mode.
SDF annotation
// SDF excerpt
(CELL (CELLTYPE "SDFFRQX1") (INSTANCE u_core/r_state_reg_3_)
(DELAY (ABSOLUTE
(IOPATH (posedge CK) Q (0.061:0.078:0.102) (0.058:0.074:0.097))))
(TIMINGCHECK
(SETUPHOLD (posedge D) (posedge CK) (0.031:0.040:0.055) (-0.012:-0.008:-0.004))
(SETUPHOLD (posedge SE) (posedge CK) (0.045:0.058:0.079) (-0.010:-0.006:-0.002))
(RECREM (posedge RN) (posedge CK) (0.020::0.035) (0.015::0.028))))
(CELL (CELLTYPE "top") (INSTANCE)
(DELAY (ABSOLUTE (INTERCONNECT u_a/Q u_b/SI (0.004:0.006:0.009)))))
// testbench
initial $sdf_annotate("top_ffg.sdf", dut, , "sdf.log", "MINIMUM");
- Triples
(min:typ:max). Pick one with the annotate call or simulator switch. Use the SDF written for the corner you care about: max delays from a slow corner for setup-type checks, min delays from a fast corner for hold.- IOPATH / INTERCONNECT
- Cell arc delays and net delays. They fill the
specifypath delays in the cell models. - TIMINGCHECK
- SETUP, HOLD, SETUPHOLD, RECOVERY, REMOVAL, RECREM, WIDTH, PERIOD. They fill
$setuphold,$recrem,$widthetc. - Negative checks
- Negative setup or hold values (common in modern libraries, as above) need the simulator to create delayed internal copies of the signals. Enable negative timing check support (e.g. VCS
+neg_tchk; check your simulator's equivalent) or the checks get clamped to 0 and you see false violations. - Annotation log
- Read it every time. Instance-name mismatches (escaped names, hierarchy separators, a netlist that doesn't match the SDF) leave cells unannotated and silently fast.
Timing checks, notifiers and X
- Legitimate violations: synchronizer first stages and async CDC flops violate by design. Disable checks on those instances with the simulator's timing-check control file, not globally.
- +notimingcheck disables all checks: fine for zero-delay functional runs, wrong for timing sign-off runs.
- X-pessimism: gate-level X propagation can be pessimistic (reconvergent X), or optimistic in RTL-style models. If a mismatch is X in simulation but ATPG expected a value, check whether the X is real (uninitialized) or a model artifact.
Serial vs parallel-load simulation
write_patterns tb_par.v -verilog -parallel and ... -serial (typically with -begin/-end to limit the serial set). The parallel testbench forces scan cell states through hierarchical references, so it's sensitive to netlist naming; rebuild it for each netlist version.Parallel load bypasses the shift path. Chain problems (broken stitching, lockup latches missing, shift hold races) are invisible until you run serial patterns. Always run the chain test pattern set serially with SDF at the fast corner.
A repeatable mismatch debug method
- Classify: how many patterns fail, which chains/cells, is it every pattern from pattern 0, from pattern N, or scattered? Pattern 0 (or chain test) failing = setup or chain problem. All cells on one chain = shift path. Scattered cells in capture = timing or logic.
- Zero-delay first: does it fail without SDF? If yes, it's functional (setup sequence, model, X), not timing.
- Map the mismatch: pattern + cycle + scan-out pin → chain + cell → instance. (Parallel testbenches usually print the instance directly.)
- Go to the capture of that pattern: dump waves around the capture window, look at the failing cell's D at the capture edge, and compare to ATPG's expected (Tessent can report simulated values per pattern for comparison).
- Trace back the cone to the first node that differs from ATPG's value or goes X. Check its clock: right number of pulses? Glitch? Check timing check messages at that time.
- Decide: real timing violation (STA gap → fix constraints/design), setup/procedure mismatch (fix test procedure or ATPG setup), model mismatch (fix library), or known issue (mask).
| Symptom | Likely cause |
|---|---|
| Chain test fails serially, one bit short | Shift hold race at a domain crossing; missing/wrong lockup latch |
| Chain test output all X | Test mode not entered (IJTAG setup), SE/clock not reaching, reset X |
| Zero-delay passes, SDF min fails in capture | Hold violation in capture path, often across clock domains or on test-mode muxed clocks |
| At-speed patterns fail, stuck-at pass | OCC not programmed as ATPG assumed, false/MCP not masked, setup at max corner |
| First pattern after setup fails only | Initialization sequence, PLL lock wait too short, state left from test setup |
| Many random cells mismatch with X | X source spreading: memory, non-scan flop, black box, uninitialized TDR |
Interview questions
If STA is clean, why run SDF GLS on scan patterns at all? core
STA is only as good as its constraints and modes. GLS checks the actual pattern sequence: test-mode setup, clock muxing and OCC pulses, SE transitions, async resets in test, and paths STA was told to ignore. It also validates the testbench and pattern formatting before tester time. It's a cross-check on the constraints, not a replacement for STA.
Parallel sim passes, serial sim fails at the first pattern. Where do you look? core
The shift path, since parallel skips it. Run the chain test serially, find the first cell where the 0011 pattern breaks, and check that transition in waves: clock arrival at both cells (skew), a lockup latch at domain crossings, Q→SI hold timing, SE value during shift, and any reset activity during shift.
Which SDF corner for which check? basic
Hold-type problems show at fast process/high voltage/low or high temperature (min delays): run min SDF from the fast corner. Setup-type at slow corner (max delays). For scan, shift is slow so hold dominates: the fast-corner run with chain patterns is the most valuable. At-speed capture needs the slow corner for setup at functional frequency.
What is a notifier and why does it matter? basic
A reg passed to a timing check system task. When the check fails, the simulator toggles the notifier; the cell's UDP treats a notifier change as "output becomes X". That's how timing violations become visible functionally in GLS.
A handful of patterns mismatch at one flop only at the slow corner at-speed. Walk me through it. deep
Likely a real setup path at functional frequency. Get the launch and capture flops and the path from the waveform, compare against STA at the same corner in at-speed capture mode. If STA shows it failing, it's a design/timing closure issue. If STA says it passes, check whether it's a mode issue (STA ran functional mode but the test-mode clock path differs), whether it's a declared multicycle/false path that ATPG should have masked (exceptions not imported to ATPG), or an SDF/annotation mismatch.
Post-silicon scan debug
Bringing scan up on first silicon, reading chain-test signatures, localizing chain and logic failures, and reading voltage/frequency behavior. This is the part where your day-to-day bring-up experience is the answer.
60-second answer
Climb in order: power/clocks/reset, TAP IDCODE, IJTAG read-back, chain integrity in bypass, then through compression, then stuck-at, then at-speed, then corners. A chain test tells you the fault type from its signature: constant means stuck, shortened runs mean slow-to-rise or slow-to-fall, a one-cycle-early stream means a hold race. Chain diagnosis finds the location using capture patterns, because cells downstream of the defect unload correctly and cells upstream don't. For logic fails, the tester fail log maps to pattern/chain/cell, and diagnosis produces ranked suspects for physical failure analysis. Shmoo shapes separate speed problems (a diagonal wall) from frequency-independent problems (a flat floor), and that tells you whether to look at setup paths or at hold/min-V issues.
Bring-up order
- Before silicon: have the chain test, a small stuck-at set and IJTAG setup vectors simulated with SDF at fast and slow corners, in the exact tester format (STIL/WGL), including test-mode entry. Prepare bypass-mode patterns and per-chain patterns.
- Tester setup is a real failure source: timing sets (strobe too early), levels, pin mapping, pattern conversion. Loopback/golden-unit checks rule it out.
- Keep a "known-good" reference die once one passes: compare failing parts against it on the same tester and program.
Chain integrity test
Chain test = load a pattern with SE held at 1 the whole time (no capture), and compare what comes out. It tests the shift path, SE and the scan clocks. Run per chain (bypass mode) to isolate. Also run at different shift frequencies and voltages: a failure that disappears at slow shift is setup/IR-drop; one that persists at any frequency is hold/race or a hard defect.
Chain diagnosis: where is the break?
Consider a chain with a stuck-at-0 at cell k (counting from scan-in). Anything that must shift through cell k reads 0. Now use a capture pattern:
- Load (the values upstream of k are destroyed by the defect, so ATPG uses patterns whose care values don't depend on those cells, or accepts X for them).
- Pulse capture: every cell captures from logic, including cells downstream of k (between k and scan-out).
- Unload: cells downstream of k reach scan-out without passing through k, so their captured values come out correctly. Cells upstream of k all pass through k and come out 0.
The boundary between "matches expected" and "all 0" in the unload stream locates k. Diagnosis tools automate this with dedicated chain diagnosis patterns, handle intermittent and timing (hold/setup) chain faults, and report a ranked list of suspect cells. Hardware assists: chain test at different frequencies and voltages, reading cell values through alternate paths (e.g. other modes), and on the bench laser voltage probing / LVI to see the toggling stop.
Logic failures and diagnosis
- Fail log
- The tester records failing cycles (vector number + pin), usually limited by fail memory depth. Convert cycle → pattern → shift position → chain → cell using the pattern file and chain map. With compression, a fail is on a compactor output: use compression-aware diagnosis or rerun in bypass.
- Diagnosis
- Tool (e.g. Tessent Diagnosis) fault-simulates candidate faults and scores how well each explains the failing and passing patterns. Output: suspects (net/cell/pin, fault type like stuck, bridge, open, cell-internal) with scores.
- Layout-aware
- Adds physical info: bridges only between neighbors, opens on specific via/segments. Narrows the area for PFA.
- Volume diagnosis
- Diagnose many failing dies and look for systematic patterns (same cell type, same layer, same layout pattern). Feeds yield learning.
- PFA
- Physical failure analysis at the suspect: delayering, SEM/TEM. Expensive, so diagnosis resolution matters.
Shmoo and voltage corners
- Vmin / LVCC: the lowest voltage that passes at a given frequency. Read it on the wall of the speed-limited plot.
- HVCC: high-voltage stress and screening; also where hold races can get worse (fast data vs clock skew) and where leakage-sensitive defects appear.
- Frequency independence is the key diagnostic: slowing the clock fixes setup problems but not hold problems.
- Temperature: with temperature inversion at advanced nodes, low temperature can be the slow corner at low voltage. Check both hot and cold.
- Guardband: production test limits sit inside the characterized pass region (voltage and frequency margin) to cover tester accuracy, aging and variation.
Triage table
| Observation | Suspect first |
|---|---|
| All chains fail, all dies | Test-mode entry, IJTAG setup, SE or scan clock at top, pad direction, tester timing/pin map |
| All chains fail, some dies | Global defect (power, clock), or marginal setup (PLL lock, reset) |
| A group of chains fails | Shared resource: clock branch, power domain, SE buffer tree, one EDT block |
| One chain, one die | Local defect → chain diagnosis |
| One chain, many dies, same cell | Systematic: hold race at a crossing, missing lockup latch, marginal cell |
| Bypass passes, EDT fails | Decompressor/compactor, channel mapping, mask logic, EDT clock, setup bits |
| Chain passes slow, fails fast shift | Shift setup or shift IR-drop; lower fshift or stagger |
| Stuck-at passes, at-speed fails | OCC setup, capture IR-drop, real slow paths, exceptions not masked |
| Fails only at low V, any frequency | Hold/race, min-V cell weakness, level shifters, retention |
| Intermittent | Noise, marginal timing, tester contact; repeat and correlate with conditions |
Interview questions
First silicon: every chain returns all 1s. What do you do? core
All chains, same signature → something global. Check the tester first (TDO/scan-out strobe, pin map, levels) on a loopback or known pattern. Then JTAG: does IDCODE read? Then the test-mode setup: read back the IJTAG/TDR bits that set scan mode, EDT bypass and OCC. Constant 1 on every output also fits scan-out pads not enabled as outputs (floating/pulled up), SE stuck, or scan clock not reaching. Work down the bring-up ladder, don't jump to chain diagnosis.
The chain output is the expected stream shifted one bit early. Diagnosis? core
A hold/race failure: somewhere a bit passes through two cells in one shift cycle, so the chain acts one cell short. Classic causes: clock skew between domains without a lockup latch, a missing/ineffective lockup latch, a negedge→posedge ordering error, or a very short Q→SI path with late capture clock. It won't go away at lower shift frequency. Try higher/lower voltage to see sensitivity, and locate it with chain diagnosis.
How does the tester fail log get mapped to a flop? basic
The tester reports failing vector (cycle) numbers per pin. The pattern file tells you which pattern and which shift cycle that vector belongs to. The pin tells you the scan-out (chain or compactor channel). Shift position + chain length gives the cell index: in unload the cell nearest scan-out comes out first. The chain map (from ATPG) gives the instance. With compression, the channel output combines several internal chains, so you need the compactor model or a bypass rerun.
At-speed yield is low but failures don't correlate to any single path. What could it be? deep
Capture IR-drop (too much switching at launch; check failing patterns' toggle activity), overtesting (non-functional paths or exceptions not masked, especially with LOS), OCC pulse width/duty problems at high frequency, or timing model/corner mismatch. Experiments: raise voltage (IR-drop and speed fails improve), use power-aware patterns, lower frequency by steps and see where fails disappear, compare against functional Fmax.
Why run chain test at more than one shift frequency? basic
To separate setup-type from hold-type chain failures. If it passes when slowed down, it's setup or shift IR-drop. If it fails at every frequency, it's a hold race or a hard defect (stuck/open).
Further reading: Advantest, An Introduction to Scan Test for Test Engineers, part 2 · Adolfsson, On Scan Chain Diagnosis for Intermittent Faults
Scan physical design, DFT timing and SDC
What DFT does to placement, routing, clock trees and timing closure, and how test modes are constrained. Round-2 material, but the basics (shift hold, modes, case analysis) come up in round 1 too.
60-second answer
Physically, DFT adds wiring that doesn't follow the functional dataflow. Scan chains get reordered after placement using the scan DEF. SE is a high-fanout net that needs its own tree, or pipelining for at-speed. Compression logic concentrates hundreds of chain ends, which causes local congestion. Test clock muxes and OCCs sit at the clock roots and affect CTS. For timing, test adds modes: shift (slow clock, SE=1, hold-critical), stuck-at capture (slow), at-speed capture (functional frequency, SE=0, exceptions matter) and JTAG/IJTAG on TCK. Each gets case analysis on test-mode and SE signals, its own clocks, and IO constraints on scan ports. The usual convergence problems are thousands of shift hold violations, SE timing, clock-mux-related paths, and missing or over-broad test-mode constraints.
Physical aspects
| Item | PD concern | Typical handling |
|---|---|---|
| Scan chains | Wirelength, congestion, hold | Scan DEF (SCANDEF) with stitching partitions and ordered segments; P&R reorders after placement and again after CTS (clock-aware) |
| Scan enable | Very high fanout; at-speed for LOS | Buffer tree, dead cycles / multicycle for LOC; pipelined SE flops per region for LOS |
| EDT decompressor/compactor | Hundreds of chain ends converge → local congestion | Place near the chains they serve; one EDT per partition; limit chains per EDT; pipeline channels |
| OCC / test clock mux | Sits at clock roots, affects insertion delay and skew between test and functional clocks | Place near PLL / clock root; balance through CTS; define clocks at OCC outputs |
| Wrapper cells | On every core port | Place near ports; shared cells avoid extra muxes on critical I/O |
| IJTAG/TAP network | Long, slow, spread over the die | Low frequency, but hold still matters; retiming stages on long hops |
| MBIST | Controllers and collars near memories | Collars at memory pins; controllers shared per cluster |
| Test points, lockup latches | Small cells on functional/scan paths | Avoid critical paths; keep lockup latch next to its launch cell |
Timing modes
| Mode | Clocks | Case analysis | What matters |
|---|---|---|---|
| Functional | PLL clocks | test_mode = 0 | Everything, as usual |
| Shift | Shift clock(s) at 10–100 MHz, through OCC slow path | test_mode = 1, SE = 1 | Hold on Q→SI (fast corner), SE is static, scan IO delays |
| Capture, stuck-at | Slow clock | test_mode = 1, SE = 0 | Mostly hold; setup trivially met |
| Capture, at-speed | PLL clocks via OCC fast path | test_mode = 1, SE = 0 | Setup at functional frequency, same exceptions as function (and ATPG must mask them), test-only clock mux paths |
| LOS (if used) | PLL via OCC | SE toggling | SE → all scan flops within one functional period |
| JTAG / IJTAG | TCK (e.g. 10–50 MHz) | per setup | TAP/TDR/SIB network setup and hold, TDI/TDO/TMS IO |
| MBIST | Functional clocks | MBIST enable | Memory interface paths at speed through collars |
SDC by mode
# ---- shift mode ------------------------------------------------
set_case_analysis 1 [get_ports test_mode]
set_case_analysis 1 [get_ports scan_en]
create_clock -name shift_clk -period 20 [get_ports shift_clk] ;# 50 MHz
set_input_delay -clock shift_clk 5 [get_ports scan_in*]
set_output_delay -clock shift_clk 5 [get_ports scan_out*]
set_false_path -from [get_ports scan_en] ;# static in this mode (it's a case value)
# ---- at-speed capture mode -------------------------------------
set_case_analysis 1 [get_ports test_mode]
set_case_analysis 0 [get_ports scan_en]
create_clock -name pll_clk -period 1.0 [get_pins u_pll/CLKOUT]
create_generated_clock -name occ_core_clk -source [get_pins u_pll/CLKOUT] \
-divide_by 1 [get_pins u_occ_core/clk_out]
# reuse functional exceptions (false paths, multicycles) for this domain
set_multicycle_path 2 -setup -from [get_clocks occ_core_clk] -to [get_clocks occ_core_clk] -through [get_pins u_div/*]
# SE is quasi-static in LOC: it changes only outside the capture window
set_false_path -from [get_ports scan_en]
# ---- JTAG ------------------------------------------------------
create_clock -name tck -period 50 [get_ports tck]
set_input_delay -clock tck 10 [get_ports {tms tdi}]
set_output_delay -clock tck 10 -clock_fall [get_ports tdo] ;# TDO launched on falling TCK
set_clock_groups -asynchronous -group tck -group [get_clocks pll_clk]
set_false_path -from scan_en is fine in shift and LOC because SE is a case value or quasi-static there. It's wrong for LOS, where SE must be timed at speed to every scan flop. Also check that constraints don't false-path test-only paths that actually toggle in the capture window, such as OCC control flops or clock-chain bits.- Mode merging
- MMMC setups often merge compatible modes to cut run time: at-speed capture is close to functional mode, but shift must stay separate (SE=1 case changes which paths exist, and clocks differ). If you merge, use
set_clock_groups -physically_exclusiveor-logically_exclusivefor clocks sharing a net. - Case analysis correctness
- One wrong
set_case_analysiscan hide a whole class of paths. Review test-mode case values against the actual DFT setup (IJTAG TDR values, OCC settings) that ATPG and GLS use. - Handoff
- Timing exceptions from SDC should feed ATPG (so it masks false/MCP endpoints at speed) and ATPG's clocking assumptions should match STA's test modes. Mismatches here are what cause "passes STA, fails tester at speed".
DFT timing convergence: the usual problems
- Shift hold: thousands of small violations after CTS. Reorder clock-aware, fix with hold buffers at SI pins, and review large-skew crossings for lockup latches.
- Test clock mux / OCC paths: glitch-free mux internal timing, clock-chain flop timing, and paths between slow and fast clocks inside OCC need proper exceptions, not blanket false paths.
- SE: transition/fanout violations on the SE tree; for LOS, real setup/hold at speed.
- Scan IO: scan-in/out through pads at shift frequency, with pad delays at corners.
- Test-mode-only paths: bypass muxes (memory bypass, clock bypass), test points, wrapper cells. These exist only in test and are easy to leave unconstrained.
- Signoff coverage: make sure every test mode is in the signoff MMMC matrix at every corner where it's hold/setup critical.
Interview questions
Why do shift paths have so many hold violations and so few setup violations? basic
Shift runs slowly, so setup has a huge margin. But Q→SI paths are very short (neighboring cells, no logic), so hold slack is roughly clock-to-Q + wire minus skew minus hold time. Any skew toward a late capture clock makes it negative, and there are as many such paths as scan cells.
What is a scan DEF and what goes in it? core
A DEF SCANCHAINS section written by the DFT tool that describes each chain as start/stop points and a set of reorderable cells, grouped into partitions, plus ORDERED segments that must stay fixed (e.g. a cell and its lockup latch, wrapper chains, cells with fixed order requirements). P&R reorders within these constraints and can swap cells between chains in the same partition. The DFT tool then needs the updated chain order back for ATPG.
How would you constrain scan_en for LOC vs LOS? core
LOC: SE changes only outside the capture window with dead cycles, so it can be treated as quasi-static: case analysis per mode, or false/multicycle from the SE port. LOS: SE must reach every scan flop between launch and capture, so it's timed as a real at-speed signal from its launch flop (a pipelined SE flop per region) with the functional clock period.
After P&R, ATPG patterns fail simulation although nothing changed in DFT. What happened? core
Likely scan reordering: chain order and possibly chain membership changed, and ATPG is still using the pre-reorder chain definitions. Re-extract the chains from the final netlist (or read back the updated scan DEF), rerun DRC/ATPG, regenerate patterns. Also hold fixing may have inserted buffers, but that doesn't change order.
Which test modes do you sign off, and at which corners? deep
Shift: hold at all fast corners, setup at slow corner at the chosen shift frequency. Stuck-at capture: hold (fast corners). At-speed capture: setup at slow corners at functional frequency, and hold at fast. JTAG/IJTAG: setup and hold at TCK frequency. MBIST: functional frequency. Plus IO timing for scan ports at the tester's timing. Mention the power/IR-drop aware analysis for at-speed capture if the team does it.
EDT deep dive: signals, interface and behavior
What the EDT block looks like from the outside, what each control signal does cycle by cycle, how a compressed pattern is actually applied, and what breaks when any of it is wrong. Tessent names are used because they're the common vocabulary; the concepts carry over to other compression schemes.
60-second answer: "walk me through the EDT interface"
From outside, EDT is a block between a few scan channels and many short internal chains. Its inputs are the channel inputs, edt_clock, edt_update, edt_bypass and scan enable (plus an optional configuration select). Its outputs are the channel outputs. Inside there's a decompressor (a ring generator plus a phase shifter), the internal chains with lockup cells where clocks can skew, and on the output side mask logic and an XOR compactor. Every pattern starts with one EDT clock pulse while edt_update is high: that resets the ring generator and moves the mask bits that were shifted in with the previous pattern into the mask hold register. Then a few initialization cycles fill the ring generator, and the shift cycles load the new pattern while the previous response unloads through the masked compactor. In capture the EDT clock is held off so the EDT state isn't disturbed. edt_bypass turns the whole thing into plain concatenated chains for debug.
The interface
edt_clock, a separate clock that pulses together with the scan shift clock and stays off in capture. edt_update is sampled on the one EDT clock pulse before each load: it resets the ring generator and copies the mask shift register into the mask hold register. edt_bypass switches the muxes that concatenate internal chains straight onto the channels. Lockup cells sit wherever the EDT clock and the chain clocks can be skewed.| Signal | Driven by | When it's active | What it does | If it's wrong |
|---|---|---|---|---|
edt_channels_in[n] | Scan-in pins (via pads, pipelines, or an SSN host) | Every shift cycle | Compressed stimulus into the ring generator | Swapped or shifted channels: every compressed pattern fails, bypass may pass |
edt_channels_out[n] | Compactor → scan-out pins | Every shift cycle | Compacted responses | Pipeline count mismatch: fails shifted by N cycles |
edt_clock | Its own pin, or a clock derived with the shift clock (OCC / SSH in hierarchical flows) | Pulses with the shift clock; off in capture | Clocks the decompressor, mask registers, pipeline flops | Pulsing in capture corrupts EDT state; not pulsing on update means no reset |
edt_update | Pin or control logic | High for the one EDT clock pulse before each load | Resets the ring generator; mask shift reg → mask hold reg | Stuck low: decompressor never re-seeded, patterns fail from the second one on |
edt_bypass | Pin or IJTAG/TDR bit (static per pattern set) | Whole pattern set | Selects bypass: internal chains concatenated onto channels | Wrong value vs. pattern set: everything fails |
scan_en | Pin, pipeline, or SSH | High in shift, low in capture | Scan cells' SE; often also qualifies mask/compactor behavior | See the test-mode timing page |
| Configuration select (optional) | Pin or TDR bit | Whole pattern set | Chooses between compression configurations (e.g. high and low compression), if the block was built with more than one | Patterns generated for the other configuration |
edt_clock, edt_update and scan enable locally. The behavior is the same; only who drives the signals changes. Say that explicitly: it shows you know the signals are logical roles, not pins.A compressed pattern, cycle by cycle
edt_update=1 resets the ring generator and moves the previous mask into the mask hold register; the chain clocks stay off in that cycle. Next come the initialization cycles: the ring generator fills up (⌈decompressor size / channels⌉ cycles). Then L shift cycles load pattern k while pattern k−1's response unloads through the compactor, masked by the mask hold register. For capture, SE falls, the EDT clock stays off, and the capture clock fires. Gaps elide most of the L shift cycles.- Update cycle. SE is high,
edt_updateis high, oneedt_clockpulse, chain clocks off. The ring generator resets to a known state and the mask bits collected during the previous load move into the mask hold register. - Initialization cycles. ⌈decompressor size ÷ channels⌉ cycles of channel data fill the ring generator before its outputs mean anything (e.g. a 64-bit ring with 8 channels → 8 cycles). The chains shift during these cycles too; the data that enters them falls off the far end by the time the load finishes.
- Shift cycles. L cycles, L = the longest internal chain. The EDT clock and the chain clocks pulse together. Each cycle, the decompressor turns n channel bits into one bit per internal chain; the compactor XORs the masked chain outputs into n channel outputs.
- Capture. SE falls, the EDT clock stays off, one or more capture pulses (OCC). EDT state is preserved.
- Repeat. The next update pulse starts pattern k+1, whose shift cycles also unload pattern k's response.
test cycles ≈ Npat × (1 update + init + L + capture cycles)
How ATPG uses this: each care bit in the chains is a linear (XOR) combination of channel bits injected so far. ATPG builds those equations for the care bits of a pattern and solves them; if they can't all be satisfied, some target faults move to another pattern. That's why dense patterns and high compression ratios cost pattern count.
Masking
- Why mask: an X in any chain feeding a compactor XOR makes that output X for that cycle, hiding every other chain on it. Masking forces the X chain's output to 0 for the pattern.
- How: mask bits are part of the compressed data. They're shifted into the mask shift register during the load, transferred to the mask hold register by the next update pulse, decoded (typically one chain, a group, or none masked), and applied at the compactor inputs during the next unload.
- Cost: a masked chain is unobservable for that whole pattern, so heavy masking means more patterns and coverage loss. Fix X at the source when you can.
- Debug use: "one-hot" masking (observe one chain at a time) turns the compactor into a transparent path for one chain. Useful for diagnosis without going to bypass.
Clocking and timing inside EDT
- Two clock domains that must behave as one in shift: the EDT clock (decompressor, masks, pipelines) and the chain clocks. They're pulsed together, but they reach flops through different trees. Where the EDT logic hands off to a chain, or a chain hands off to the compactor, lockup cells absorb the skew, exactly like domain crossings inside a chain.
- EDT logic is shift-mode logic: timed at shift frequency with SE=1 case analysis. Hold dominates, like chains. The decompressor and compactor have real XOR depth, so at high shift frequency setup can matter too.
- Channel pipelines: long routes between pins and EDT blocks get pipeline flops on both the input and output channels. They add latency; ATPG and the pattern files must know the count. Identical broadcast destinations need equal pipeline depth.
edt_updatetiming: it's quasi-static (it changes only between cycles with no conflicting clock), so it's usually constrained like SE: case analysis per mode, plus a check that it settles before the update pulse.- Capture: the EDT clock must be off. If the EDT clock is derived from a shared clock source, the gating must guarantee no pulse in the capture window.
Bypass and configurations
| Mode | What changes | Used for |
|---|---|---|
| Compressed (normal) | Channels → decompressor → internal chains → compactor → channels | Production patterns |
| Bypass | Muxes concatenate internal chains onto channels (with lockups); decompressor and compactor out of the path | Silicon bring-up, chain integrity, diagnosis, fallback if EDT logic is defective |
| Low-compression config (if built) | Fewer internal chains per channel or a different compactor setup | Better diagnosis resolution, pattern count at lower compression, debug |
| Low-power shift (if built) | Decompressor holds values for chains with no care bits | Shift power reduction |
What actually breaks, and how you'd see it
| Symptom | Likely cause | First check |
|---|---|---|
| Bypass chain test passes, every compressed pattern fails | EDT control: update/clock sequencing, channel mapping, configuration bits, pipeline count | Simulate one compressed pattern with waves on edt_clock/edt_update; compare the test procedure against the EDT spec |
| First compressed pattern passes, later ones fail | Ring generator not re-seeded (update not reaching), or mask transfer broken | edt_update path and its timing |
| Fails shifted by a fixed number of cycles | Pipeline stage count differs from what ATPG assumed | Pipeline count in the EDT/pattern setup vs. netlist |
| Fails only on chains next to EDT logic | Missing or wrong lockup between EDT and chain clocks (hold) | SDF GLS at the fast corner; STA shift-mode hold at the handoff |
| No fails but coverage lower than expected | Excess masking, X sources | ATPG X/mask statistics per chain |
| Diagnosis can't localize | Compactor mixes many chains per channel | Rerun in bypass or low-compression configuration |
Interview questions
What does edt_update do, and when is it asserted? core
It's sampled on one EDT clock pulse at the start of each load (in the load/unload procedure, before shifting). On that pulse the ring generator resets to a known state, and the mask shift register contents (the mask bits that came in with the previous pattern) are copied into the mask hold register, which then masks the unload of that previous pattern's response. It's low during shift and during capture.
Why is the EDT clock separate from the scan clock, and why must it be off in capture? core
The EDT logic has its own state (ring generator, mask registers, pipelines) that must advance exactly once per shift cycle and must not change during capture. A separate clock lets it be pulsed in lockstep with shift but gated off in capture, independent of whatever functional clocks capture uses. If it pulsed in capture, the ring generator and mask state would shift away from what ATPG computed.
What are initialization cycles and how many are there? core
After the reset on update, the ring generator needs channel data before its outputs carry useful, independent values. The tool adds about ⌈decompressor size ÷ number of channels⌉ extra shift cycles per load for this. They're part of every pattern's shift length, so a very wide decompressor with few channels costs noticeable test time.
Why is the mask applied one pattern later? core
Because the response of pattern k is unloaded while pattern k+1 is being loaded. The mask bits for k come in with k's load (into the mask shift register), the update pulse before load k+1 moves them into the hold register, and they gate the compactor inputs during that load, which is exactly when k's response is coming out.
Compressed patterns fail on silicon, bypass passes. Walk me through debug. deep
Bypass passing means chains, SE, shift clocks and scan pins work. So the problem is in what compression adds: the EDT control sequence (update pulse, EDT clock gating), channel pin mapping and pipeline stages, configuration bits, and the EDT logic itself. I'd simulate one compressed pattern at gate level with the exact tester procedure and look at edt_update/edt_clock around the load, check the channel order and pipeline count against what ATPG used, and check the configuration TDR values actually applied. On silicon, try the low-compression configuration or one-hot masking to see whether fails are channel-wide (control) or chain-specific (a real defect).
Where do lockup cells go in an EDT design? core
Wherever a shift path crosses from one clock to another with possible skew: between the decompressor (EDT clock) and the first cell of each internal chain, between the last cell and the compactor/pipeline flops, and in the bypass path where chains are concatenated. The tool inserts them automatically when it can see the clocks differ.
How does compression interact with X sources? basic
An X on any chain feeding an XOR in the compactor corrupts that output for that cycle, so it hides all the other chains on it. EDT uses per-pattern masking to block X chains, which costs observability. Many X sources means more masking, more patterns and lower effective compression, so X sources are fixed at the design level (initialize memories, bound black boxes, avoid non-scan flops).
How would you pick the number of channels and internal chains? core
Channels come from available pins (or SSN bandwidth). Internal chains set chain length: total cells ÷ chains. More chains per channel shortens shift but makes each pattern harder to encode and increases compactor fan-in, so pattern count rises and diagnosis gets harder; routing around the EDT block gets congested. I'd run the tool's compression analysis on a few ratios and pick the knee of test data volume vs. pattern count, then check routing.
Sources: EDT compressor and controller (mask registers, edt_update) · EDT decompressor (ring generator, phase shifter) · EDT user manual excerpts (EDT clock off in capture, initialization cycles, lockups, pipelines) · EDN, test compression
Test-mode timing: pins, scan enable and capture
The timing questions that come up once you've explained scan: how a scan bit gets from a GPIO pin into the first cell safely, why scan-out is a tester strobe problem, how you guarantee capture only happens after scan enable has settled everywhere, and which checks belong to which mode.
60-second answer: "how do you time scan from a GPIO pin?"
The reference isn't an on-chip flop, it's the tester. The tester drives scan-in at a known point in each cycle and pulses the scan clock at another point, usually mid-cycle. Data goes pad → test mux → first cell on a short path; the clock goes pad → OCC or clock mux → the whole clock tree. So at the first cell the clock arrives late relative to the data, and scan-in is mainly a hold check: the cell has to sample the current bit before the tester's next bit gets there. Scan-out is the other direction: the last cell launches on its late clock, the bit goes through a mux and an output pad, and it has to be stable at the tester's strobe, so that's a setup check against the strobe time. In STA you model the tester with a clock at the scan clock port plus input and output delays that match the tester's timing set. If a path doesn't close you add a retiming flop near the pad, add hold delay on the input, move the strobe, or slow shift down.
Scan-in from a pin
Put numbers on it. Tester period T, data driven at the start of the cycle, clock edge at the pin at T/2. The first cell sees data after tin and the clock after tclk:
hold: T/2 + tclk + thold < T + tin → margin = T/2 + tin − tclk − thold
- Insertion delay eats hold margin. Every ns of clock tree between the pin and the first cell comes straight out of the margin. Scan clocks through an OCC and a big tree can have several ns.
- Tester edge placement matters. The T/2 in the formula is the tester's choice. Moving the clock edge earlier relative to data drive (or driving data later) changes the margin. That's why test engineers and DFT agree on a timing set per pattern type.
- Slowing shift helps hold here, unlike on-chip hold, because the margin contains T/2. It's still the wrong first fix; fix the path.
- Fixes: hold buffers on the scan-in route (the setup side has lots of room), a retiming flop right after the pad clocked from a low-insertion point or on the opposite edge, or a pipeline stage between pad and EDT inputs, which ATPG then has to know about.
One way to write the SDC for this
# shift mode, T = 20 ns. Tester drives SI at 0, pulses scan_clk rising at 10. create_clock -name scan_clk -period 20 -waveform {10 15} [get_ports scan_clk] # virtual clock = the tester's drive/strobe reference create_clock -name ate -period 20 -waveform {0 10} set_input_delay -clock ate 1.0 [get_ports {scan_in*}] ;# tester + board skew set_output_delay -clock ate 3.0 [get_ports {scan_out*}] ;# strobe setup + board set_case_analysis 1 [get_ports scan_en] set_case_analysis 1 [get_ports test_mode] set_propagated_clock [all_clocks]
Tools differ and teams have their own templates. The point is that the input and output delays are relative to a clock that represents the tester, and the real scan clock is defined with its true waveform at the pin, so the tool computes the pin-to-cell relationship with the actual insertion delay.
Scan-out to the tester
strobe must land inside that window, with the tester's setup/hold around it
- It's a window, not a single number. Early strobes see the previous bit, late strobes see the next one. At slow shift the window is wide and centered wherever you like. As shift speeds up the window shrinks and moves relative to the tester edge.
- Fixes: strobe later or in the following cycle (the tester program then accounts for the one-cycle offset), a retiming flop or lockup latch near the output pad to hold the data longer, better pad drive, or lower shift frequency.
- Pipelines change the pattern, not just timing. Any retiming flop in the scan-in or scan-out path adds a cycle of latency. The pattern generation setup has to include it, otherwise every compare is shifted by one.
GPIO pad details that come up
| Item | Why it matters in test |
|---|---|
| Direction control | Scan pins are usually shared with functional GPIO. In test mode the pad's output enable is forced by test_mode so scan-in pads are inputs and scan-out pads are outputs regardless of functional logic. That forcing mux is on the path and in STA it's a case-analysis value. |
| Pad delay | IO pads with level shifters are slow compared to core logic (on the order of ns). Pad delay dominates tin and tout, and it varies a lot across corners. |
| Bidirectional turnaround | If a pad switches direction between phases (e.g. used as input in shift and output elsewhere), the switch needs its own time and must not happen while the tester drives it. |
| Output switching | Many scan-out pins toggling at once cause simultaneous switching noise on the IO supply. It's a reason to limit output count or stagger strobes. See the note on the SSN page. |
| Tester and board | Tester edge placement accuracy and board trace delay go into the input/output delay values. Test engineering usually provides them. |
Making sure capture happens after SE has settled
Scan enable is a very high-fanout net that has to change twice per pattern: fall before capture, rise before the next shift. Two conditions per transition:
- Falling edge, early side (hold): SE must not reach any cell before the last shift edge plus hold, or that cell captures functional data instead of its shift neighbor. Because SE is launched from the tester or a flop clocked by the same shift clock, this is an ordinary hold check.
- Falling edge, late side (setup): SE must reach the farthest cell and settle before the capture edge minus setup. In LOC there's plenty of time if the protocol adds dead cycles or if the OCC takes a few cycles to synchronize SE before it releases capture pulses. In LOS the launch happens on the last shift edge, so SE has only one at-speed period to fall everywhere.
- Rising edge: after capture, SE must reach every cell before the first shift edge of the next load. Usually easy at shift speed, but check it.
| Technique | What it buys |
|---|---|
| Dead cycles / extra tester cycles | Time for SE to settle before the OCC or tester releases capture. Costs a few cycles per pattern. |
| OCC synchronization | The OCC sees SE fall, synchronizes it into the fast clock domain, and only then generates the capture pulses. That delay is built-in settling time and guarantees ordering. |
| SE pipelining | Flops (clocked by the shift clock) re-drive SE per region, so each local tree is small. The pipeline adds cycles of latency that the test procedure accounts for. |
| STA treatment | For LOC, SE is not a single-cycle at-speed path: time it at shift frequency, or use a multicycle or case-analysis setup that matches the protocol. For LOS it has to be timed at speed, which is a main reason LOS is used less. |
Which checks belong to which mode
| Mode | Clocks | What dominates | Typical signoff corner |
|---|---|---|---|
| Shift | Slow scan clock, SE=1 by case analysis | Hold: chain neighbors, lockups across domains, scan-in from pins; setup at scan-out pins | Hold at fast corners; also check shift speed limit at slow corners |
| Capture (LOC, at-speed) | Functional clocks via OCC, SE=0 | Functional setup and hold, same as mission mode but only for the clock pairs pulsed together | Setup at slow corner, hold at fast |
| Capture (stuck-at, slow) | Slow clocks, SE=0 | Hold (setup relaxed) | Fast |
| SE transitions | Shift clock | Hold always; setup vs. protocol time | Both |
| JTAG / IJTAG | TCK | Its own domain; TDI/TMS sample on rising TCK, TDO changes on falling | Both, at TCK frequency |
| Static config | none | test_mode, TDR bits, EDT bypass: set by case analysis per mode; paths from them are usually false paths | n/a |
Interview questions
Why is scan-in from a pin usually a hold problem? core
The data path is short (pad, mux, a route) and the clock path is long (pad, OCC or mux, full clock tree). The tester drives the next bit a fixed time after its clock edge, so if the clock at the first cell is late by more than that time minus hold, the cell sees the next bit instead of the current one. Setup has most of a period of slack.
The scan-in hold check fails by 1 ns at the fast corner. What are your options? core
Add hold delay on the scan-in route (cheap, setup has room). Add a retiming flop close to the pad clocked from a low-insertion point, or a negative-edge retiming stage. Ask test engineering whether the timing set can drive data later relative to the clock. Check whether the input delay values are realistic. Slowing shift helps too, but that's a last resort because it costs test time.
How do you make sure capture only happens after SE has reached every flop? core
Protocol plus timing. The protocol gives SE time: dead cycles between the last shift and capture, and in OCC designs the controller synchronizes SE into the fast domain before it fires capture pulses. Timing then proves SE reaches the farthest cell within that time (at shift frequency or with a matching multicycle), and that it doesn't reach the nearest cells before the last shift edge plus hold. If the tree is too slow, pipeline SE per region.
Why is LOS harder on scan enable than LOC? core
In LOS the launch is the last shift edge and the capture follows one at-speed period later, so SE has to switch from 1 to 0 at every cell within that period. It becomes an at-speed signal with huge fanout. In LOC, SE falls during a slow gap before the launch pulse, so it has many ns to settle.
Scan-out at the tester fails only above 50 MHz shift. What's going on? core
The valid window at the pin (starting at clock pin → insertion → clk-to-Q → mux → output pad → board) slides relative to the tester strobe as the period shrinks, and at some point the strobe sees the old or the next bit. Check the output path delay vs. strobe placement. Fixes: strobe later or in the next cycle, a retiming stage near the pad, better pad drive, or keep shift at the lower frequency.
What happens to your patterns if you add a pipeline flop between the scan-in pad and the EDT channel? basic
Every bit arrives one shift cycle later. The pattern generation setup must declare the pipeline so the load data is shifted accordingly, and for outputs the expected data is shifted too. Forgetting it gives a pattern set where everything miscompares by one position.
How are static test signals like test_mode handled in STA? basic
Set by case analysis per mode, which removes the other mux leg from timing. Paths starting at those signals are usually false paths because they're set once before the test and held. SE is the exception people get wrong; it switches every pattern.
Which corner would you sign off shift timing at? core
Hold at the fast corners (fast process, high voltage; and check both temperature extremes because of temperature inversion at advanced nodes). Also check that the chosen shift frequency meets setup at the slow corner, which matters for scan-out and EDT logic with XOR depth.
SSN: Streaming Scan Network
SSN is how Tessent delivers scan data to many cores through a fixed-width, packetized bus instead of wiring each core's EDT channels to chip pins. Below: the problem it solves, how the bus and packets work, how it's configured, what it changes for timing and debug, and how to answer if "SSN" turns out to mean simultaneous switching noise.
60-second answer: "tell me about SSN"
In a hierarchical design each core has its own EDT with a certain number of channels. The classic way to test them at the top is to mux core channels onto chip scan pins, so you test a few cores at a time, you fix the channel allocation at design time, and when cores have different chain lengths the short ones sit padded. SSN replaces that with a bus of fixed width, sized by the pins you have, that runs through a Streaming Scan Host in every core. Data goes across the bus as packets: one packet holds all the bits every active core needs for one shift cycle. Each host knows its bit positions and shift count, which are programmed through IJTAG at pattern time, picks its bits out of the stream, and generates its core's scan enable, EDT controls and clock gating locally. So you can test any group of cores together without changing hardware, cores can shift independently and capture together, identical cores can get the same data with the comparison done on chip, and the bus width is independent of the cores' channel counts.
The problem it solves
- Pins don't scale with cores. A chip with dozens of cores, each with several EDT channels, can't bring every channel to a pin. Pin-muxed schemes test cores in groups, and the grouping is frozen in hardware.
- Padding waste. When cores tested together have different chain lengths or pattern counts, the shorter ones wait. The tester is busy but those cores aren't getting useful data.
- Late changes are expensive. If a core's pattern count grows late in the project, the best grouping changes, but the top-level muxing is already built.
- Scan enable and capture must be coordinated globally when many cores share pins, which makes top-level timing harder.
The bus and the hosts
Packets
- Packet size = sum of the channel bits every active host needs for one internal shift cycle. It's independent of bus width.
- No wasted bus bits. Packets don't need to align with bus cycles, so an odd total doesn't leave idle lanes.
- Independent shift. A core with a shorter chain can finish its load and wait, or be given fewer bits per packet; hosts manage their own shift and capture timing within the pattern.
- Shift rate per core ≈ bus bits per second ÷ packet size. Adding active cores makes packets bigger, so each core shifts slower, but the tester pins are always fully used.
SSN vs. pin-muxed channels
| Pin-muxed EDT channels | SSN | |
|---|---|---|
| Top-level wiring | Per-core channels routed to pin muxes | One bus through all hosts |
| Which cores test together | Fixed by hardware grouping | Chosen at pattern generation |
| Channels per core | Limited by pins at the top | Set by core needs; bus width set by pins |
| Different chain lengths | Padding | Bits sized per core, no padding |
| Scan enable, EDT controls | Top-level signals to every core | Generated in each host |
| Identical cores | Broadcast possible, outputs need pins or compaction | Broadcast plus on-chip compare, sticky status |
| Setup cost | Low: simple muxes | IJTAG configuration per pattern set, host area, more complex debug |
Published results from Siemens and users report large test time reductions (Siemens cites up to several times on some designs, and Intel has reported double-digit percentage cycle reductions). Quote numbers carefully in an interview and frame them as "published results", not guarantees.
Timing and debug implications
- The bus is a timing path like any other. Hop to hop between hosts at the bus frequency; long hops get pipeline stages. Identical cores that share broadcast data should see equal pipeline depth.
- Local scan enable. Because each host generates its own core's SE and EDT controls, the huge chip-wide SE tree goes away. The SE timing problem from the timing page becomes a per-core problem, clocked locally.
- Clock crossing at the host. The host bridges the bus clock and the core's shift clock. Treat it like any other domain crossing in shift mode: lockups or designed handshakes, and constraints from the tool's generated SDC.
- Debug path. Bring up the IJTAG network first, then the SSN bus in a simple configuration (one host active, known data), then each core's EDT in bypass through its host, then compressed patterns. Same ladder as always, one layer more.
- Diagnosis. Fails come back as bus positions and cycles; the tool maps them back to core, channel and pattern. With on-chip compare you get pass/fail per host, and you rerun the failing core in a streaming (not compare) setup for full fail data.
If they meant simultaneous switching noise
SSN is also the old abbreviation for simultaneous switching noise, and it's relevant to test. It's fine to ask "streaming scan network or switching noise?" and then answer. The short version:
- Core side: during shift, many flops toggle every cycle; during capture, a large fraction of the logic switches at once. That draws current spikes and causes IR drop and ground bounce, which slow paths (false at-speed fails) or corrupt shift. Mitigations: low-power fill (fill don't-care bits to reduce toggling), shift at lower frequency, Q-gating, staggered or partitioned capture clocks, capture power limits in ATPG. See the low-power scan page.
- IO side: many output pads switching together (scan-out pins, boundary scan EXTEST) cause bounce on the IO supply. Mitigations: limit simultaneous outputs, stagger, proper power/ground pad ratios.
Interview questions
Why would you use SSN instead of muxing core channels to pins? core
Because the core grouping and channel allocation become pattern-time decisions instead of hardware decisions, the bus width is set by pins independent of how many channels the cores have, cores with different chain lengths don't pad each other, SE and EDT controls are generated locally in each core, and identical cores can be tested in parallel with on-chip compare. It matters most with many cores and limited pins.
What is a packet in SSN? core
All the bits that every active core needs for one internal shift cycle. If core A has 5 channels and core B has 4, a packet is 9 bits. Packets stream across the bus regardless of bus width, so on an 8-bit bus the packet boundaries don't line up with bus cycles. Each host knows its offset in the packet and the packet size.
How is each host told what to do? core
Through IJTAG before the pattern set: the host's registers hold whether it's active, its position in the packet, the packet size and shift count, and related control. The pattern generation tool writes this setup as part of the test procedure.
How does SSN test identical cores efficiently? core
The same stimulus goes to all copies, with the expected response also sent on the bus. Each host compares its core's output locally and records a sticky pass/fail. So adding copies doesn't add bus bits. After the pattern set, the per-host status is read out. If a core fails, you rerun it alone to collect full fail data for diagnosis.
What does SSN change about scan enable timing? deep
The chip-level SE tree is gone. Each host generates SE for its core, timed in the core's shift domain, so the SE setup/hold problem shrinks to core size and is closed at core level. Cores can even be in capture while others are still shifting, as long as the tool schedules it that way.
SSN patterns fail on all cores, but IJTAG reads and writes work. Where do you look? deep
The bus itself: pin mapping of ssn_in/ssn_out, bus pipeline count vs. what the tool assumed, bus clock and any frequency conversion settings, and whether the host configuration actually applied (read back the host registers). Then run the simplest configuration: one host active, low bus rate. If that passes, add hosts until it breaks.
"SSN" in the sense of switching noise: why does it matter for test? basic
Test switches much more logic at once than functional mode: shift toggles many flops per cycle and capture can switch a large part of the design at once. The resulting IR drop and ground bounce slow paths and can cause fails that aren't real defects, or corrupt shift. You control it with low-toggle fill, lower shift frequency, partitioned capture and ATPG power constraints.
Sources: Siemens white paper, Tessent Streaming Scan Network · SemiEngineering, SSN: an efficient packetized data network · 2020 ITC paper (Siemens blog) · On-chip compare and diagnosis with SSN
Scripting drills
The scripts a DFT engineer actually writes: read a file line by line, pull out what you need with a regex, count or group it in a dictionary, print a clean report. No computer-science tricks needed. Each problem opens to plain-language steps, then a hint, pseudo-code and a short tested Python solution.
What the interviewer is actually looking for
- It works on the example. Correct output on the sample beats clever code.
- You understand the file. Say what a line looks like and what you're pulling out of it. This is where your hardware background is the advantage.
- It's readable. Clear names, one small function per job, a comment where the format is odd.
- You think about messy input. Comments, blank lines, a line that doesn't match, a name that appears twice. Mention them even if you only handle some.
- You talk while you write. A slightly wrong script you explained well scores better than a silent perfect one.
Nobody expects algorithm theory in a hardware scripting round. If they ask "how would this scale to a 5 GB netlist?", the answer is practical: read line by line instead of all at once, and don't redo the same work twice.
The Python you actually need
Almost every problem below uses only these pieces.
Read a file line by line
for line in open("run.log"):
line = line.strip() # remove spaces and the newline at the end
if not line or line.startswith("#"):
continue # skip blank lines and commentsLine numbers
for line_no, line in enumerate(open("run.log"), start=1):
...Split a line into words
parts = "chain1 u_core/a_reg_0_".split() # ['chain1', 'u_core/a_reg_0_']
chain, cell = partsRegex: find a pattern and pull out pieces
import re
m = re.search(r"pattern=(\d+)\s+chain=(\S+)", line)
if m:
pattern = int(m.group(1)) # first ( ) group
chain = m.group(2) # second ( ) group
# \d digit \w letter/digit/_ \S non-space \s space + one or more * zero or more . any charCount things with a dictionary
counts = {}
for word in ["C4", "S1", "C4"]:
counts[word] = counts.get(word, 0) + 1 # .get returns 0 the first time
# counts == {"C4": 2, "S1": 1}Group things: dictionary of lists
files_of = {}
files_of.setdefault("core", []).append("rtl/core.v") # creates the list the first timeSort a report
for name in sorted(counts): # by name
...
for name in sorted(counts, key=lambda k: counts[k], reverse=True): # biggest count first
...Compare two lists: sets
before = {"a", "b", "c"}
after = {"b", "c", "d"}
before - after # {'a'} removed
after - before # {'d'} added
before & after # {'b', 'c'} in bothWalk a folder tree
import os
for folder, subfolders, files in os.walk("rtl"):
for name in files:
if name.endswith(".v"):
path = os.path.join(folder, name)Nice printing
print(f"{name:12s} {count:4d} {slack:+.3f}") # padded name, 4-wide int, signed 3 decimalsCommand-line arguments
import sys
log_path = sys.argv[1] # python3 summary.py run.logA function that calls itself (recursion)
def read_flist(path, files):
for line in open(path):
if line.startswith("-f"):
read_flist(line.split()[1], files) # a nested list: handle it exactly the same way
else:
files.append(line.strip())
# Use it when the data nests inside itself: file lists that include file lists,
# modules that contain modules. Each call handles one level.File formats you'll be handed
Simulator file list (flist)
// top.f
+incdir+rtl/include // where `include files live
+define+SYNTH // a macro
rtl/top.v // a source file
rtl/core.v
-f rtl/sub/sub.f // another file list, read it too
-v $LIBDIR/cells.v // library file
rtl/top.v // duplicate: keep the first one
-f means "read this other list too". Lines starting with + or - are options, not files. Comments are // or # depending on the tool.
RTL vs post-synthesis netlist
module core #(parameter W = 8) ( // RTL: directions inside ( )
input clk,
input [W-1:0] a,
output [W-1:0] y
);
alu #(.W(W)) u_alu (.a(a), .y(y)); // instance: TYPE #(params) NAME ( ... );
always @(posedge clk) ... // behavior, no gates yet
endmodulemodule core_W8_0 ( clk, a, y, si, se ); // netlist: names only in ( )
input clk, si, se; // directions declared below
input [7:0] a;
output [7:0] y;
wire n12, n13;
alu_W8_1 u_alu ( .a(a), .y({n12, n13}) ); // child module, renamed by synthesis
SDFFRQX1 \y_reg[0] ( .D(n12), .SI(si), .SE(se), .CK(clk), .Q(y[0]) );
AOI22X1 U137 ( .A0(a[2]), .A1(n12), .B0(n13), .B1(y[0]), .Y(n14) );
endmodule| RTL | Netlist | |
|---|---|---|
| Port directions | Inside the ( ) header | Names in the header, input/output lines below |
| Contents | always, assign, expressions | Almost only instances of library cells |
| Names | Written by people | Tool-made: U137, n12, y_reg_3_, alu_W8_1 |
| One instance | Usually one line | Often several lines: split on ;, not on newlines |
| Escaped names | Rare | \y_reg[0] is one name: starts at \, ends at the next space |
;. Each piece is one statement, however many lines it took. Look at its first word to decide what it is (input, wire, assign, or a cell/module type).Warm-ups
P1Summarize errors and warnings in a tool log basic
Given a DFT tool log, print how many times each error/warning ID appears and the line where each first appears.
Input looks like
// Tessent-style log excerpt (made up)
Info: Reading netlist top.v
Warning: Clock 'clk_div' is gated by non-test logic. (C4)
Warning: Clock 'clk_aux' is gated by non-test logic. (C4)
Error: Scan cell u_core/r_reg_3_ has uncontrollable reset. (S1)
Warning: Clock 'clk_div' is gated by non-test logic. (C4)
Info: 12,480 scan cells found
Error: Scan cell u_core/q_reg_7_ has uncontrollable reset. (S1)
Warning: Non-scan cell u_pll/lock_q is an X source. (D5)Think it through
- Loop over the lines with line numbers.
- Keep only lines that start with
Error:orWarning:and end with an ID in parentheses. - Build a key like
Warning C4; add 1 to its count. - If it's the first time you see that key, remember the line number.
- Print sorted.
Hint
One regex does the matching and pulls out both pieces: (Error|Warning):.*\((\w+)\). Two dictionaries: one for counts, one for first line.
Pseudo-code
for each line (with its number):
if it looks like "Error: ... (ID)" or "Warning: ... (ID)":
key = kind + ID
counts[key] += 1
if key not seen before: first_line[key] = line number
print each key, its count and first linePython solution
import re
def log_summary(path):
"""Count errors and warnings by message ID, remember where each first appeared."""
counts = {} # "Warning C4" -> how many times
first_line = {} # "Warning C4" -> first line number where it appeared
for line_no, line in enumerate(open(path), start=1):
m = re.match(r"(Error|Warning):.*\((\w+)\)\s*$", line)
if not m:
continue
key = m.group(1) + " " + m.group(2) # e.g. "Warning C4"
counts[key] = counts.get(key, 0) + 1
if key not in first_line:
first_line[key] = line_no
for key in sorted(counts):
print(f"{key:12s} x{counts[key]:<3d} first at line {first_line[key]}")
return countsOutput on the sample:
Error S1 x2 first at line 5
Warning C4 x3 first at line 3
Warning D5 x1 first at line 9Follow-ups they might ask
- Also print the text of the first message for each ID (store the line, not just its number).
- Exit with a non-zero code if there's any error, so a regression script can stop (
sys.exit(1)). - Ignore a list of waived IDs read from a file.
P2Find every line that uses a signal basic
Search every .v/.sv file under a folder and print file:line: text for each line containing a given name, like a small grep.
Input looks like
rtl/
top.v
core.v
sub/alu.v
search word: DFFX1Think it through
- Walk the folder tree.
- For each Verilog file, loop over its lines with line numbers.
- If the word is in the line, print file, number and the line.
Hint
os.walk gives you every folder and its files. if word in line is enough; a regex is only needed for whole-word matching.
Pseudo-code
for each folder, files under root:
for each file ending in .v or .sv:
for each line (with number):
if word in line: print path:number: linePython solution
import os
def grep(root, word, exts=(".v", ".sv")):
"""Print file:line: text for every line that contains word, in all Verilog files under root."""
hits = 0
for folder, subfolders, files in os.walk(root):
for name in sorted(files):
if not name.endswith(exts):
continue
path = os.path.join(folder, name)
for line_no, line in enumerate(open(path), start=1):
if word in line:
print(f"{path}:{line_no}: {line.strip()}")
hits += 1
return hitsOutput on the sample:
rtl/core.v:3: DFFX1 r0 (.D(a[0]), .CK(clk), .Q(y[0]));
rtl/sub/alu.v:2: DFFX1 f0 (.D(a[0]));
rtl/sub/alu.v:3: DFFX1 f1 (.D(a[1]));Follow-ups they might ask
- Whole-word only, so
clkdoesn't matchclk_div:re.search(r"\b" + word + r"\b", line). - Skip comment lines.
- Only search files listed in an flist (combine with P3).
File lists and Verilog
P3Expand a file list (flist) core
Return the list of source files in an flist, including the files from any nested -f lists, without duplicates, in order.
Input looks like
// top.f
+incdir+rtl/include // where `include files live
+define+SYNTH // a macro
rtl/top.v // a source file
rtl/core.v
-f rtl/sub/sub.f // another file list, read it too
-v $LIBDIR/cells.v // library file
rtl/top.v // duplicate: keep the first oneThink it through
- Read the flist line by line; drop comments and blank lines.
- If the line is
-f other.f, read that file the same way and add its files to the same list. - If the line starts with
+or-, it's an option: skip it (or collect it separately). - Otherwise it's a file: add it unless it's already in the list.
Hint
The nested -f is the only tricky part. Write one function that reads one flist, and when it sees -f, it calls itself on the nested file. That's all "recursion" means here.
Pseudo-code
function read_flist(path, files):
for each line in path:
remove comments; skip if empty
if line starts with "-f": read_flist(nested path, files)
else if line starts with + or -: skip (option)
else if line not in files: add it
return filesPython solution
import os
def read_flist(path, files=None):
"""Return the list of source files in a file list, following nested -f files.
Paths are taken relative to where the script runs (the -f convention)."""
if files is None:
files = [] # first call: start an empty list
for line in open(path):
line = line.split("//")[0].split("#")[0].strip() # drop comments
if not line:
continue
parts = line.split()
if parts[0] == "-f": # nested file list: read it the same way
read_flist(parts[1], files) # <- the function calls itself
elif line.startswith("+") or line.startswith("-"):
continue # +incdir+, +define+, -v, -y ... skipped here
elif line not in files: # keep first occurrence only
files.append(line)
return filesOutput on the sample:
['rtl/top.v', 'rtl/core.v', 'rtl/sub/alu.v']Follow-ups they might ask
- Also collect
+incdir+directories and+define+macros into their own lists. - Replace
$VARwith the environment variable (os.path.expandvars). -F: paths inside the nested list are relative to that list's own folder (os.path.dirname+os.path.join).- Detect a list that includes itself (keep a list of flists you're currently reading).
P4Which file defines each module? basic
Given a set of Verilog files, list every module with the file(s) it's defined in, and flag modules defined more than once.
Input looks like
rtl/top.v: module top ( ... ); ... endmodule
rtl/core.v: module core ( ... ); ... endmodule
rtl/core_old.v: module core ( ... ); ... endmodule <- duplicateThink it through
- Read each file and remove comments (so a commented-out
module foodoesn't count). - Find every
module NAME. - Add the file to that name's list.
- Print; mark names with more than one file.
Hint
re.findall(r"\bmodule\s+(\w+)", text) returns all module names in the text. The \b keeps it from matching inside endmodule.
Pseudo-code
for each file:
text = file contents without comments
for each "module NAME" in text:
where[NAME].append(file)
for each NAME: print NAME and its files, flag if more than onePython solution
import re
def remove_comments(text):
text = re.sub(r"/\*.*?\*/", "", text, flags=re.S) # /* ... */ blocks, can span lines
return re.sub(r"//.*", "", text) # // to end of line
def module_map(paths):
"""Return {module_name: [files that define it]}."""
where = {}
for path in paths:
text = remove_comments(open(path).read())
for name in re.findall(r"\bmodule\s+(\w+)", text):
where.setdefault(name, []).append(path)
return where
def print_report(where):
for name in sorted(where):
files = where[name]
flag = " <-- defined more than once!" if len(files) > 1 else ""
print(f"{name:12s} {', '.join(files)}{flag}")Output on the sample:
alu rtl/sub/alu.v
core rtl/core.v, rtl/core_copy.v <-- defined more than once!
pll ip/pll/pll.v
top rtl/top.vFollow-ups they might ask
- Which one does the simulator use? The one compiled first/last depending on the tool; say so and suggest failing the flow on duplicates.
- Also find modules that are used but never defined (compare with the instance list from P5).
P5Find all instances of a module or cell core
Given Verilog files and a top module, print the full hierarchical path of every instance of module (or cell) X, e.g. every DFFX1 under top.
Input looks like
module top(...); core u_core (...); core u_core2 (...); pll u_pll (...); endmodule
module core(...); alu u_alu (...); DFFX1 r0 (...); endmodule
module alu(...); DFFX1 f0 (...); DFFX1 f1 (...); endmoduleThink it through
- First build a table: for each module, the list of (child type, instance name) inside it.
- An instance statement looks like
TYPE NAME ( ... );orTYPE #( ... ) NAME ( ... );. Skip keywords likeinput,wire,assign. - Then start at the top. For each child: if its type is X, record the path. If the child is itself one of your modules, go inside it and repeat, with the path extended.
Hint
Two separate steps make this easy: (1) parse every module into a table once, (2) walk the table from the top. The walk is the same "function calls itself" pattern as P3.
Pseudo-code
table = {module: [(child_type, inst_name), ...]} # step 1
function find(module, path): # step 2
for each (child_type, inst_name) in table[module]:
child_path = path + "/" + inst_name
if child_type == X: record child_path
if child_type is in table: find(child_type, child_path)
find("top", "top")Python solution
import re
KEYWORDS = {"module", "input", "output", "inout", "wire", "reg", "assign", "always",
"if", "else", "begin", "end", "for", "case", "parameter", "localparam"}
def remove_comments(text):
text = re.sub(r"/\*.*?\*/", "", text, flags=re.S)
return re.sub(r"//.*", "", text)
def read_modules(paths):
"""Return {module_name: [(child_type, instance_name), ...]} from Verilog files.
Handles 'type name (' and 'type #(params) name (' on one statement."""
modules = {}
for path in paths:
text = remove_comments(open(path).read())
for m in re.finditer(r"\bmodule\s+(\w+)(.*?)\bendmodule", text, flags=re.S):
name, body = m.group(1), m.group(2)
children = []
for statement in body.split(";"):
inst = re.match(r"\s*(\w+)\s*(?:#\s*\(.*\))?\s*(\\\S+|\w+)\s*(\[[^\]]*\])?\s*\(",
statement, flags=re.S)
if inst and inst.group(1) not in KEYWORDS:
children.append((inst.group(1), inst.group(2).lstrip("\\")))
modules[name] = children
return modules
def find_instances(modules, top, target, path=None, found=None):
"""Every hierarchical path under top whose instance type is target."""
if path is None:
path, found = top, []
for child_type, inst_name in modules.get(top, []):
child_path = path + "/" + inst_name
if child_type == target:
found.append(child_path)
if child_type in modules: # it's one of our modules: look inside it too
find_instances(modules, child_type, target, child_path, found)
return foundOutput on the sample:
top/u_core/u_alu/f0
top/u_core/u_alu/f1
top/u_core/r0
top/u_core2/u_alu/f0
top/u_core2/u_alu/f1
top/u_core2/r0Follow-ups they might ask
- Only count them: return a number instead of paths. For a huge design, count per module once and multiply, instead of walking every copy.
- Match by pattern, e.g. every cell starting with
SDFF:child_type.startswith("SDFF"). - Instance arrays
u_alu[3:0]: say there are 4 copies and multiply.
P6Print a module's ports with direction and width basic
Given one module's text, print each port with its direction and width. It must work for RTL-style and netlist-style headers.
Input looks like
// RTL style
module m (input wire [3:0] a, b, output reg y);
// netlist style
module alu_W8_0 ( clk, a, y );
input clk;
input [7:0] a;
output [7:0] y;Think it through
- Grab the module name, the text inside the header
( ), and the body. - If the header contains
input/output, it's RTL style: go item by item; an item without a direction keeps the previous one (bis alsoinput [3:0]). - Otherwise it's netlist style: the header only has names; read the
input ...;/output ...;statements in the body.
Hint
Notice the two styles out loud before coding. That alone shows you've looked at real netlists.
Pseudo-code
name, header, body = regex on the module
if header has "input"/"output":
for each comma-separated item: take direction/range if present, else reuse previous
else:
for each statement in body: if it starts with input/output, record direction/range for each name
print ports in header orderPython solution
import re
def port_list(text):
"""Return (module_name, [(port, direction, range)]) for the first module in text.
Works for ANSI headers (RTL) and non-ANSI headers (netlists)."""
text = re.sub(r"/\*.*?\*/", "", text, flags=re.S)
text = re.sub(r"//.*", "", text)
m = re.search(r"\bmodule\s+(\w+)\s*(?:#\s*\(.*?\)\s*)?\((.*?)\)\s*;(.*?)\bendmodule", text, flags=re.S)
name, header, body = m.group(1), m.group(2), m.group(3)
decl = r"(input|output|inout)\s+(?:wire\s+|reg\s+|logic\s+)?(\[[^\]]*\])?\s*(.*)"
ports = []
if re.search(r"\b(input|output|inout)\b", header): # ANSI: directions are in the header
direction, rng = None, ""
for item in header.split(","):
d = re.match(decl, item.strip(), flags=re.S)
if d: # new direction on this item
direction, rng, pname = d.group(1), d.group(2) or "", d.group(3)
else: # same direction as the item before
pname = item
ports.append((pname.strip(), direction, rng))
else: # non-ANSI: names only in the header
info = {}
for statement in body.split(";"):
d = re.match(decl, statement.strip(), flags=re.S)
if d:
for pname in d.group(3).split(","):
info[pname.strip()] = (d.group(1), d.group(2) or "")
for pname in header.split(","):
pname = pname.strip()
if pname:
ports.append((pname,) + info.get(pname, ("?", "")))
return name, portsOutput on the sample:
('core', [('clk', 'input', ''), ('a', 'input', '[W-1:0]'), ('y', 'output', '[W-1:0]')])
('alu_W8_0', [('clk', 'input', ''), ('rst_n', 'input', ''), ('a', 'input', '[7:0]'), ...])Follow-ups they might ask
- Compare RTL and netlist port lists of the same block: which ports did scan insertion add?
- Convert a range like
[7:0]into a width of 8.
Netlists
P7Count cells and find flops that aren't scan flops core
From a gate-level netlist, count instances of each cell type and list the flops that have no scan-in pin (a quick check after scan insertion).
Input looks like
SDFFRQX1 \y_reg[0] ( .D(n1), .SI(si), .SE(scan_en), .CK(clk), .RN(rst_n), .Q(y[0]) );
DFFRQX1 cnt_reg ( .D(n3), .CK(clk), .RN(rst_n), .Q(n4) );
AND2X1 U12 ( .A(a[0]), .B(b[0]), .Y(n1) );Think it through
- Split the netlist on
;so each instance is one piece, even across lines. - Skip pieces that start with
module,input,wire,assign... - For the rest: first word = cell type, second = instance name (strip a leading
\), then pull every.PIN(net)into a small dictionary. - Counting: add 1 per cell type.
- Non-scan flops: it has
CKandDpins but noSIpin.
Hint
Store each instance's pins as a dictionary ({"D": "n1", "SI": "si", ...}). Then questions become simple checks like "SI" in pins, and P8 reuses the same parsing.
Pseudo-code
for each statement in netlist split on ";":
skip declarations
type, name, pins = parse it
counts[type] += 1
if pins has CK and D but no SI: non_scan.append(name)Python solution
import re
def read_cells(path):
"""Return a list of (cell_type, instance_name, {pin: net}) for every instance in a netlist."""
text = open(path).read()
text = re.sub(r"//.*", "", text)
cells = []
for statement in text.split(";"): # one statement can span many lines
words = statement.split()
if not words or words[0] in ("module", "endmodule", "input", "output", "inout",
"wire", "assign", "reg"):
continue
m = re.match(r"\s*(\w+)\s+(\\\S+|\w+)\s*\((.*)\)\s*$", statement, flags=re.S)
if not m:
continue
cell_type, inst = m.group(1), m.group(2).lstrip("\\")
pins = dict(re.findall(r"\.(\w+)\s*\(\s*([^()]*?)\s*\)", m.group(3)))
cells.append((cell_type, inst, pins))
return cells
def cell_counts(cells):
counts = {}
for cell_type, inst, pins in cells:
counts[cell_type] = counts.get(cell_type, 0) + 1
return counts
def non_scan_flops(cells):
"""Flops without a scan-in pin. Checking the SI pin is more reliable than the cell name."""
result = []
for cell_type, inst, pins in cells:
is_flop = "CK" in pins and "D" in pins
if is_flop and "SI" not in pins:
result.append(inst)
return resultOutput on the sample:
{'SDFFRQX1': 2, 'DFFRQX1': 1, 'AND2X1': 1, 'XOR2X1': 1, 'INVX1': 1, 'BUFX2': 1, 'alu_W8_0': 1, 'LATNQX1': 1}
['cnt_reg']Follow-ups they might ask
- Why checking the
SIpin beats checking the nameSDFF*: libraries name cells differently. - Some non-scan flops are on purpose (synchronizers). Read an allowed list from a file and skip those.
- Counts per module rather than for the whole file: track the current
modulename while you loop.
P8Trace a scan chain through the netlist core
Starting from the scan-in net, follow the chain: find the flop whose SI is on that net, go to its Q, and repeat. Pass through buffers. Print the chain in order and the net where it ends.
Input looks like
SDFFRQX1 \y_reg[0] ( ..., .SI(si), ..., .Q(y[0]) );
SDFFRQX1 \y_reg[1] ( ..., .SI(y[0]), ..., .Q(y[1]) );
BUFX2 U15 ( .A(y[1]), .Y(so) );Think it through
- Using P7's parsing, build a lookup: net → the scan flop whose SI is on it. Also net → buffer whose input is on it.
- Start with
net = scan_in. - Loop: if a flop's SI is on this net, record it and move to its Q net. If a buffer's input is on it, move to its output. Otherwise stop: this is the end of the chain.
- Add a safety limit so a wiring loop can't make it run forever.
Hint
The lookup dictionary is the key: "which flop listens on this net?" is then one line instead of searching the whole netlist every step.
Pseudo-code
si_of = {net: flop whose SI is on net}
net = scan_in
while True:
if net in si_of: record the flop; net = its Q
elif a buffer reads net: net = buffer output
else: stop
print the order and the final net (should be a scan-out port)Python solution
def trace_chain(cells, scan_in_net):
"""Follow a scan chain: find the cell whose SI is on the current net, move to its Q, repeat.
cells: list of (cell_type, instance, {pin: net}) as returned by read_cells()."""
si_of = {} # net -> scan cell whose SI pin is on that net
buf_in = {} # net -> buffer whose input is on that net (chains often pass through buffers)
for cell_type, inst, pins in cells:
if "SI" in pins:
si_of[pins["SI"]] = (inst, pins)
elif cell_type.startswith("BUF"):
buf_in[pins.get("A")] = (inst, pins)
order = []
net = scan_in_net
while True:
if net in si_of:
inst, pins = si_of[net]
order.append(inst)
net = pins["Q"]
elif net in buf_in:
inst, pins = buf_in[net]
net = pins["Y"] # step through the buffer, don't count it
else:
break # nothing continues: this net is the chain's end
if len(order) > 1_000_000:
raise RuntimeError("chain loops back on itself?")
return order, netOutput on the sample:
(['y_reg[0]', 'y_reg[1]'], 'so')Follow-ups they might ask
- Compare the traced order with the chain report from the DFT tool (P11 does the comparison).
- Two flops listening on the same net = a branch: report it instead of picking one.
- The chain ends on something that isn't a scan-out port: that's a broken chain, print where.
Tester logs and reports
P9Fail log to instance names core
Count tester fails per scan cell and print the worst cells first, with their instance names from the chain map.
Input looks like
# tester: pattern=... lines
pattern=12 chain=c3 cell=40 expected=1 actual=0
pattern=13 chain=c3 cell=40 expected=1 actual=0
pattern=13 chain=c1 cell=7 expected=0 actual=1
pattern=20 chain=c3 cell=40 expected=1 actual=0
# chain map: <chain> <cell> <instance>
c3 40 top/u_core/u_alu_2/f1
c1 7 top/u_pll/lock_qThink it through
- Read the chain map into a dictionary: (chain, cell) → instance.
- Read the fail log; for each fail line add 1 to that (chain, cell).
- Sort by count, largest first, and print with the instance name (or say it's missing from the map).
Hint
Use a pair (chain, cell_number) as the dictionary key. Convert the cell number with int() in both files so they match.
Pseudo-code
cmap = {(chain, cell): instance} from the chain map
for each fail line: fails[(chain, cell)] += 1
for key in fails, biggest first: print count, key, cmap.get(key, "missing")Python solution
import re
def read_chain_map(path):
"""chain map lines: '<chain> <cell_index> <instance>' -> {(chain, index): instance}"""
cmap = {}
for line in open(path):
parts = line.split()
if len(parts) == 3 and not line.startswith("#"):
cmap[(parts[0], int(parts[1]))] = parts[2]
return cmap
def fail_summary(log_path, cmap):
"""Count fails per (chain, cell); print worst first with the instance name."""
fails = {}
for line in open(log_path):
m = re.search(r"pattern=(\d+)\s+chain=(\S+)\s+cell=(\d+)\s+expected=(\d)\s+actual=(\d)", line)
if not m:
continue
key = (m.group(2), int(m.group(3)))
fails[key] = fails.get(key, 0) + 1
for key in sorted(fails, key=lambda k: fails[k], reverse=True):
inst = cmap.get(key, "<not in chain map>")
print(f"{fails[key]:4d} fails {key[0]}[{key[1]}] {inst}")
return failsOutput on the sample:
3 fails c3[40] top/u_core/u_alu_2/f1
1 fails c1[7] top/u_pll/lock_qFollow-ups they might ask
- Also show whether it's always 1→0 or 0→1 (a stuck-at hint).
- How many different patterns failed on each cell?
- Given only tester cycle numbers, convert them to pattern and cell first (P12).
P10Timing report summary basic
From a timing report with many paths, print the worst slack (WNS), the sum of negative slacks (TNS) and the number of failing paths for each path group.
Input looks like
Startpoint: u_core/reg_a (rising edge-triggered flip-flop clocked by clk_core)
Endpoint: u_core/reg_b (rising edge-triggered flip-flop clocked by clk_core)
Path Group: clk_core
Path Type: max
slack (VIOLATED) -0.042
Startpoint: u_core/reg_c (rising edge-triggered flip-flop clocked by clk_core)
Endpoint: u_core/reg_d (rising edge-triggered flip-flop clocked by clk_core)
Path Group: clk_core
Path Type: min
slack (MET) 0.013
Startpoint: tck_in (input port clocked by tck)
Endpoint: u_tap/ir_reg_0 (rising edge-triggered flip-flop clocked by tck)
Path Group: tck
Path Type: max
slack (VIOLATED) -0.200Think it through
- Each path is a block of lines. Remember the current
Path Groupwhen you see it. - When you reach the
slackline, the block is complete: take the number. - If it's negative, update that group's worst, total and count.
Hint
You don't need to understand the whole report: just two kinds of lines, the group name and the slack. Keep "current group" in a variable between lines.
Pseudo-code
group = None
for each line:
if "Path Group:" in line: group = the name
if it's a slack line:
slack = number
if slack < 0: wns[group] = min(...); tns[group] += slack; count[group] += 1
print per groupPython solution
import re
def timing_summary(path):
"""Worst slack (WNS), total negative slack (TNS) and number of failing paths per path group."""
results = {} # group -> [wns, tns, count]
group = None
for line in open(path):
m = re.search(r"Path Group:\s*(\S+)", line)
if m:
group = m.group(1) # remember which group this path belongs to
m = re.search(r"slack\s*\(\w+\)\s*(-?[\d.]+)", line)
if m and group:
slack = float(m.group(1))
if group not in results:
results[group] = [0.0, 0.0, 0]
if slack < 0:
results[group][0] = min(results[group][0], slack)
results[group][1] += slack
results[group][2] += 1
for group, (wns, tns, n) in results.items():
print(f"{group:10s} WNS {wns:+.3f} TNS {tns:+.3f} failing {n}")
return resultsOutput on the sample:
clk_core WNS -0.042 TNS -0.042 failing 1
tck WNS -0.200 TNS -0.200 failing 1Follow-ups they might ask
- Split setup (
Path Type: max) and hold (min) results. - List the 5 worst endpoints.
- Group by launch/capture clock pair to spot a bad clock crossing.
P11What changed between two chain reports? basic
Compare the scan chain report before and after place-and-route: which cells moved to a different chain, which disappeared, which are new.
Input looks like
before: after:
chain1 u_core/a_reg_0_ chain1 u_core/a_reg_0_
chain1 u_core/a_reg_1_ chain1 u_core/b_reg_0_
chain1 u_core/a_reg_2_ chain1 u_core/a_reg_2_
chain2 u_core/b_reg_0_ chain2 u_core/a_reg_1_
chain2 u_core/b_reg_1_ chain2 u_core/b_reg_1_
chain2 u_core/spare_reg chain2 u_core/new_regThink it through
- Read each report into a dictionary: cell → (chain, position).
- For every cell in before: missing in after → removed; different chain → moved.
- For every cell in after that wasn't in before → added.
Hint
Think of it as two dictionaries keyed by cell name. Once both are built, the comparison is a few ifs.
Pseudo-code
before = {cell: chain}, after = {cell: chain}
for cell in before:
if cell not in after: removed
elif chains differ: moved
for cell in after not in before: addedPython solution
def read_chains(path):
"""'<chain> <cell>' per line, in scan order -> {cell: (chain, position)}"""
where = {}
position = {}
for line in open(path):
parts = line.split()
if len(parts) != 2:
continue
chain, cell = parts
position[chain] = position.get(chain, -1) + 1
where[cell] = (chain, position[chain])
return where
def chain_diff(before_path, after_path):
before, after = read_chains(before_path), read_chains(after_path)
for cell in sorted(before):
if cell not in after:
print(f"removed : {cell} (was {before[cell][0]})")
elif before[cell][0] != after[cell][0]:
print(f"moved : {cell} {before[cell][0]} -> {after[cell][0]}")
for cell in sorted(after):
if cell not in before:
print(f"added : {cell} ({after[cell][0]})")Output on the sample:
moved : u_core/a_reg_1_ chain1 -> chain2
moved : u_core/b_reg_0_ chain2 -> chain1
removed : u_core/spare_reg (was chain2)
added : u_core/new_reg (chain2)Follow-ups they might ask
- Also report chain lengths before/after (count per chain).
- Report cells that stayed in the same chain but changed position (reorder).
- Why this matters: ATPG must use the chain order from the final netlist, or patterns fail.
Bits and numbers
P12Bit tricks and cycle-to-cell math basic
Small functions that come up in hardware scripting: count 1s, parity, get/set/clear a bit, reverse bits, hex to a binary string, and convert a tester fail cycle into (pattern, scan cell).
Input looks like
count_ones(0b1011) -> 3
reverse_bits(0b0011, 4) -> 0b1100 (12)
hex_to_bits('0x3A', 8) -> '00111010'
cycle_to_cell(1202, setup=1000, chain_length=100) -> (2, 99)Think it through
- Bits: shifting (
>>,<<) and masking (& 1,|,& ~) cover almost everything. - Cycle to cell: subtract the setup cycles; each pattern takes chain_length + capture cycles; the remainder is the shift position; the first bit out is the cell nearest scan-out, so cell = chain_length − 1 − position.
Hint
Python helpers: bin(x) gives '0b1011', int('3A', 16) parses hex, format(x, '08b') gives 8 binary digits.
Pseudo-code
count_ones(x): count "1" characters in bin(x)
get_bit(x, n): (x >> n) & 1
set_bit(x, n): x | (1 << n)
clear_bit(x, n): x & ~(1 << n)
cycle_to_cell: offset = cycle - setup
pattern = offset // (L + capture)
position = offset % (L + capture)
cell = L - 1 - positionPython solution
def count_ones(value):
"""Number of 1 bits, e.g. to count toggles or check parity."""
return bin(value).count("1")
def parity(value):
return count_ones(value) % 2 # 1 = odd number of ones
def get_bit(value, n):
return (value >> n) & 1
def set_bit(value, n):
return value | (1 << n)
def clear_bit(value, n):
return value & ~(1 << n)
def reverse_bits(value, width):
result = 0
for i in range(width):
if get_bit(value, i):
result = set_bit(result, width - 1 - i)
return result
def hex_to_bits(hex_string, width):
"""'0x3A', 8 -> '00111010' (MSB first)"""
return format(int(hex_string, 16), f"0{width}b")
def cycle_to_cell(cycle, setup_cycles, chain_length, capture_cycles=1):
"""Tester fail cycle -> (pattern, cell index). Assumes each pattern = shift of chain_length + capture.
The first bit out of the chain comes from the cell nearest scan-out."""
offset = cycle - setup_cycles
per_pattern = chain_length + capture_cycles
pattern = offset // per_pattern
shift_pos = offset % per_pattern
cell = chain_length - 1 - shift_pos
return pattern, cellOutput on the sample:
3 1 12 00111010
(2, 99)
(2, 0)Follow-ups they might ask
- Parity of a 32-bit word without
bin(): loop over bits and XOR them. - Which bits differ between expected and actual?
expected ^ actual, then list the set bits. - Chains of different lengths: the shift length is the longest chain, shorter chains are padded at the start.
Stretch: only if you have time
These come up less often. Knowing the idea in plain words is enough; you'd write them the same way as the problems above.
- Fan-in cone
- "Which flops and inputs feed this flop's D pin?" Build a dictionary net → the cell that drives it (P7 parsing, using output pin names like
Q,Y). Start at the D pin's net, look up its driver, then look at that driver's input nets, and keep going backwards. Stop at flops and module inputs. Keep a list of nets already visited so you don't repeat. - Balancing scan chains
- "Split these blocks into N chains so the longest is as short as possible." Simple approach: sort blocks biggest first, and give each one to whichever chain is currently shortest. It's not always perfect, but it's close and it's what you'd do by hand.
- Combinational loop check
- "Is there a loop of gates with no flop in it?" Follow gate outputs to gate inputs, skipping flops; if you ever get back to a gate you're currently in the middle of exploring, there's a loop. Mention it; it's rarely asked to be coded in full.