Before the Cluster: Reconciling the Predictors and Their Documents
Abstract
The branch prediction cluster (BPC) instantiates the Pacino predictors and provides the interface to the fetch target queue (FTQ). These sessions covered three preconditions necessary before the BPC integration could be performed.
During the statistical corrector (SC) planning phase the requirements from the TAGE to SC interface had changed requiring a TAGE fix to match the new package(s).
While planning for the integration I realized the planning documents had been built in isolation and needed a correlation effort to find inconsistent statements and any misleading/incomplete information.
Before writing the FTQ interface I decided there needed to be a frontend
planning document that encompassed how the BPC and FTQ would inter-operate.
These decisions are contained in fe_decisions.md.
BPC integration was not performed in these sessions but all three of the preconditions were completed. At this stage the interfaces within the FE were not fully described. That occurred in the next set of sessions.
The primary methodology finding in these sessions was reinforcement of the role of the architect in the design process. Document analysis reported the RTL and planning documents had drifted, either through mistakes or missing requirements, in a few cases with no single source to defer to. In the Pacino methodology these decisions are made by the human.
Reconciling TAGE with SC package changes
In sessions 057 and 058 edits to bp_structs_pkg.sv were made as part of SC
planning [1]. cond_pred_meta_t, cond_pred_upd_inp_t and tage_high_conf
were retired.
BP-081 retyped the TAGE update queue and response buffer to the existing
tage_upd_inp_t and tage_pred_meta_t, and deleted the dead tage_high_conf
logic, closing TD #94 and TD #95.
BP-081 also generated the SC-facing fields described in TD #87/#88. These were
the one-hot strong, medium and weak decode of the provider counter and the
extended counter 2*ctr - 7.
BP-081’s only package edit was the re-addition of tage_pred_weak for this new
decode. The tage_pred_strong definition changed from “not weak” to an
explicit strong encoding, 000 or 111. As part of this the UAON[F1] gate moved
from the old !tage_pred_strong to tage_pred_weak, 011 or 100.
tb_tage_tasks.sv used the removed field but was missing from the task
manifest. The IA stopped and asked before editing outside its scope, and I
authorized the fix. BUG-006 now requires affected files to be found by
searching the unit, not taken from a list in the task.
I asked the IA if it had run all targets in the bpu Makefile. It had run make
all but not all targets were covered. The missing targets were run with no
errors. TD #99 was recorded for future work to develop a pre-push regression
scheme.[F2]
The coverage run reported 73.7% line coverage for cov_tage and 79.5% for
cov_tage_table, against an earlier stated figure above 90%. TD #100 was
recorded for diagnosis.[F3]
Session 060 ended with all seven predictors and bp_history passing at the
unit level.
Auditing the planning documents
Prior to BPC integration I redirected session 061 to audit the documents. The audit was three read-only IA tasks, FTB, then TAGE/SC and finally ITTAGE/RAS. Each task compared the planning documents for the group against the group’s testbenches, packages, RTL, Makefile targets, as well as shared files. The task files reported discrepancies and any information that PA/I thought might be misleading. The IA audit results were reported to the PA and PA drafted the planning document updates. A discrepancy that needed a design decision rather than a text correction came to me.
Findings by group
INFRA-008 found no discrepancies in the FTB documents, and two stale RTL comments (TD #104).
INFRA-009 found 12 discrepancies in TAGE and SC. tage_pred_strong was
redefined to 000 or 111, from “not weak”. The TAGE base table index is PC bits
12 to 2, not 11 to 1. br_imli_mode is now an RTL parameter instead of a
port. The SC threshold is dynamic, not fixed, and the SC table geometry had
changed. The SC unit testbench instantiates only ST1, so ST0 has no unit
coverage.
INFRA-010 found 10 discrepancies in ITTAGE [4] and RAS. The ITTAGE
allocation-write field order, MSB to LSB, is TAG, TGT, EPC, USE, CTR, VALID,
incorrectly documented as TAG, EPC, USE, CTR, TGT, VALID. The ITTAGE table tag
width is IT_TBL_TAG, 8 to 11 bits per table, not as documented 38b. Two target
write strobes were present in the RTL, prm_tgt_wr_u0 and alt_tgt_wr_u0, not
a single strobe. The parameter ITTAGE_RESP_BUF_DEPTH was marked as
vestigial; the response buffer is no longer in the design.
The remaining minor findings were stale names and descriptions, and an unread RAS input deferred to TD #101.
The documents were corrected for all findings. The two design discrepancies are covered in the next section.
Two rulings
Two ITTAGE findings were contradictions between documents, detected and reported by the IA. There was confusion as to where indirect calls were assigned; the documents differed. I made the ruling that ITTAGE predicts the target, and the RAS manages the return address. This is conventional.
The second contradiction was the description of the longest ITTAGE table, IT5. A cut and paste error labeled IT5 as a BrIMLI table [5], a transposition from the SC documents.
This gap was recorded as TD #102 for later resolution. TD #102 will impact both the ITTAGE and the history module but at present it represents a loss of prediction accuracy rather than a functional failure.
The base table’s initial value
During audit there was some confusion on the initialization value used by the
TAGE base table (T0). All TAGE tables, T0 included, initialize from
TAGE_SRAM_INIT_VALUE, which is 0 (strongly not taken). T0 is intended to
initialize to weakly taken, 10, as the planning document stated. The audit
cleanup mistakenly changed the document to match the RTL. TD #103 now captures
the fix: a T0-specific init value of 10, and restoration of the document.
Specifying the FTQ boundary
Session 062 developed the fe_decisions.md file which defines the cluster
organization and the cluster to FTQ interface with role assignments for
exchanging predictions, redirects and updates. It settled seven points:[F4]
- The FTQ entry is dual-slot: block-level fields (PC, history pointers, RAS snapshot) and a per-slot array (target, branch type, direction, source).
- The two slots are the two branch fields of one 32-byte FTB block, from one lookup, not two fixed PC ranges.
- Predictors do not drive redirects. Each presents its prediction at its stage, and the cluster compares it with the earlier one and derives the redirect.
- ITTAGE produces its final target at p2.
- A history checkpoint is the GHR and PHR pointer pair, one per FTQ entry. The RAS snapshot is a separate field restored by the same redirect.
- An FTQ entry is allocated at p1 for every prediction block, including a p1 miss, so a later predictor always has an entry to redirect against.
- At most one RAS operation occurs per prediction block. A call or return is a taken branch, so one in slot 0 ends the block before slot 1. Each FTQ entry therefore needs only one RAS snapshot.
The interface file
The interface file was not written in these sessions. The file needed all
eight predictor port lists. Some files shared in the PA session arrived empty
according to the PA. It is likely this was due to the length of the files and
the remaining context available in the PA session. Unfortunately Claude.ai has
no /context command and it has no way to report context load, unlike Claude
Code.[F5]
The port inventory moved to the IA, which read the repository directly. In the
next session, ftq_bpu_interfaces.md was written from that inventory.
Experiment Summary
| Experiment | Description | Status | Checks | Runtime | Context |
|---|---|---|---|---|---|
| BP-081 | TAGE struct reconciliation; TD #87 confidence decode and TD #88 extended counter generated; UAON gate moved | PASS | sim_tage 105/0, sim_tage_table 15/0, make all exit 0, remaining targets run separately | 36m 29s | 34% |
| INFRA-008 | Read-only audit, FTB documents against RTL | COMPLETE | no discrepancies; 2 stale RTL comments | 3m 35s | 13% |
| INFRA-009 | Read-only audit, TAGE and SC documents against RTL | COMPLETE | 12 discrepancies | 8m 37s | 6% (main thread) |
| INFRA-010 | Read-only audit, ITTAGE and RAS documents against RTL | COMPLETE | 10 discrepancies, 2 escalated | 8m 23s | 6% |
BP-081’s run time was recorded before the follow-up run of the targets outside
make all, and its context figure after it. INFRA-009 read its roughly 50 files
in three sub-agents with separate context windows, reported at about 350k
tokens combined; the 6% is the main thread only.
Design Process Notes
The IA contribution
BP-081 went furthest beyond its task specification. The IA identified that
narrowing tage_pred_strong would change the UAON update condition, gated the
update on tage_pred_weak to preserve the prior behavior, and verified the
change against an existing directed test. It also identified that the UAON
rules document still carried the superseded definition. When it found an
affected testbench outside its manifest, it halted and requested authorization
instead of editing the file.
The audits cited file and line for each finding and traced several to their cause, including the BrIMLI description of IT5 and the split target-write strobe. INFRA-010 reported the two cross-document contradictions without picking a side, and noted that the IT5 item was marked complete on the side the RTL does not implement. INFRA-008 recorded that it had not re-run the FTB suite and why.
The PA contribution
The planning assistant wrote the four task files, drafted the document
corrections from the audit findings, and edited fe_decisions.md with me.
The BP-081 manifest was built from BP-080’s reference list instead of a search
of the unit. The same failure had been recorded one session earlier, and it is
why BUG-006 exists. The assistant grouped the T0 initial value with the text
corrections. In session 062 it asked which of two documents governed the slot
model when the documents in front of it settled the question: ftb_decisions.md
was complete, later, and named the status-file items it superseded.
My contribution
I folded the TD #87 and TD #88 generation into BP-081, asked whether every
target had run, and authorized the out-of-scope testbench edit. I redirected
session 061 to the audit, required the later audit groups to read the shared
documents in full, and made the two ITTAGE rulings. I recorded TD #103. In
session 062 I supplied the design corrections to fe_decisions.md: the
single-stage ITTAGE target, the slot model, the checkpoint definition and the
redirect model.
The generalization
An audit that compares a document with the RTL finds where they disagree. It does not find which one is wrong. An earlier case of a reference taken from the design under test is described in [6].
Twenty of the 22 findings were resolved by changing the document to match the RTL. For stale names, removed parameters and superseded ports that is correct, because the RTL is the later artifact and the change was deliberate. The two findings escalated for a ruling were those where documents disagreed with each other, so there was no single RTL value to defer to. T0’s initial value was different. It was a value, the document stated the intent, and the RTL did not implement it. The audit treated it like a stale name.
IT5 shows the same thing from the other side. One set of documents matched the RTL, and the RTL was missing logic.
For this project, a finding that changes a documented value, as opposed to a name, a width citation or a cross-reference, needs the owner to state the intended value before the document is edited toward the RTL. That is the decision a consistency audit cannot make on its own.
What comes next
The range closes with all seven predictors and bp_history passing at the unit
level, their planning documents checked against the RTL, and the FTQ boundary
written down in fe_decisions.md. The cluster has still not been started.
The next session continues the interface work that session 062 could not finish: the IA’s inventory of every port on the eight top-level modules, the FTQ-to-cluster interface file written from it, then the cluster itself.
The debt opened in this range is small RTL work that the cluster does not need in order to elaborate: the unread RAS input (TD #101), IT5’s missing folds (TD #102), T0’s initial value (TD #103) and the two FTB comments (TD #104). The SC connections that the cluster will make, TD #89 to #92, and the end-to-end fold check, TD #84, remain as the previous range left them.
Technical Debt Referenced
The table below reports status as of the close of this range. Later experiments outside the range have since changed the state of some items.
| # | Item | Resolution path |
|---|---|---|
| 87 | TAGE strong/medium/weak decode. | CLOSED BP-081. One-hot on the post-mux provider CTR: strong {000,111}, weak {011,100}, medium the rest. tage_pred_strong narrowed from NOT WEAK; the UAON update gate moved to tage_pred_weak. |
| 88 | TAGE extended counter generation. | CLOSED BP-081. tage_extd_ctr = 2*ctr - 7, signed 5b, range -7 to +7. tage_provider_ctr remains an internal signal, not a struct field. |
| 94 | bp_arb_spec reconciliation to the standalone SC. | CLOSED BP-081. The two tage.sv FIFOs retyped off the retired cond_pred_* structs. |
| 95 | tage_pred_meta_t changed; TAGE references to reconcile. | CLOSED BP-081. Dead tage_high_conf logic deleted. |
| 99 | Create a PR/CI/CD process. | Motivating evidence (session-060): make all silently omits sim_ittage, sim_tage_manual and the cov targets. CI must run every target, not make all. |
| 100 | TAGE line coverage below the previously stated >90%. | cov_tage 73.7%, cov_tage_table 79.5% (BP-081). Decide whether this is genuine under-coverage of the new decode logic or an accounting difference; add directed coverage or correct the number. |
| 101 | ras.sv declares input ras_pc_p2, unread in the module. | Undocumented before session-061; ras_interfaces.md now lists it and cites this item. Confirm whether needed at RAS cleanup; if not, remove from ras.sv and tb_ras.sv. |
| 102 | bp_history.sv does not generate IT5 folds. | ittage.sv wires it_t5_idx_fh/tag_fh1/tag_fh2 to outputs that are never driven, permanently 0. IT5 has real history (hist 32, FH 9, FH1 9, FH2 8). Add IT5 fold generation, same pattern as IT1-IT4. |
| 103 | TAGE T0 initial value. | T0 initializes to 00, strongly not-taken. Intended is 10, weakly taken. Impacts TAGE_SRAM_INIT_VALUE with sram_init. When closed, tage_cntrl_decisions.md’s T0 line must be updated again; the session-061 correction documented current behavior only. |
| 104 | Two stale FTB RTL comments (INFRA-008). | ftb_cntrl.sv states 107 bits per way where the entry is 105; ftb.sv states ftb_fastpath_en is beyond the interface draft, which now lists it. Comment-only; fold into the first FTB RTL task. |
References
[1] “The Statistical Corrector: Design Choices at p3” (BLOG_bpu_17), for the package edits that broke TAGE and the BP-080 investigation that scoped the repair.
[2] Seznec, André, and Pierre Michaud. “A case for (partially) tagged geometric history length branch prediction.” The Journal of Instruction-Level Parallelism 8 (2006): 23.
[3] Seznec, André. “Tage-sc-l branch predictors.” JILP-Championship Branch Prediction. 2014.
[4] Seznec, André. “A 64-Kbytes ITTAGE indirect branch predictor.” JWAC-2: Championship Branch Prediction. 2011.
[5] Seznec, André, Joshua San Miguel, and Jorge Albericio. “The inner most loop iteration counter: a new dimension in branch history.” Proceedings of the 48th International Symposium on Microarchitecture. 2015.
[6] “External Anchors: When a Proof and Its Reference Share the Same Error” (BLOG_bpu_16), for the earlier case of a reference taken from the design under test.
Footnotes
[F1] UAON is known in the literature as USE_ALT_ON_NA, i.e.
use-alternate-on-newly-allocated. It is a counter that tracks whether the
alternate prediction is more accurate than the primary prediction when the
primary’s entry appears newly allocated, i.e. its counter is weak (011 or
100). See [2] and [3].
[F2] TOOLS-006, in a later range, closed TD #99 with tools/regress.sh, which
runs every target of every rtl/ Makefile and is run by a pre-push hook.
[F3] BP-097, in a later range, found that 8,043 of the 8,700 lines counted by
cov_tage are testbench lines, and that the DUT-only coverage is near 90%.
TD #100 remains open until a conclusion states its denominator.
[F4] fe_decisions.md has been revised substantially since session 062. In
session 064 the RAS top-of-stack read was moved from p1 to p0, and after a
port inventory confirmed that no predictor declares a redirect port, the
redirect section was rewritten to two cluster-level groups named by stage,
bpu_redir_p2 and bpu_redir_p3. In session 071 the 32-byte unit was named
the prediction block, distinct from the 64-byte fetch block the IFU reads.
This post uses the later term.
[F5] Confirmed with the latest Claude.ai.
Q: do you have any method to return your context load, remaining, etc ?
A: No. I have no tool that reports context used or remaining, and no visibility into token counts. Claude Code shows the IA a “Ctx %”, which is where your task headers get that figure; nothing equivalent exists in this chat.
Jeff Nye is a microprocessor architect with 35 years of industry experience spanning performance modeling, RTL implementation, and architecture for high-performance OOO processors. He has contributed RTL to Pentium 4, ARM V7, TI C6x and RISC-V designs, and recently served as sole architect and full-stack implementer of the TAGE-SC-L + ITTAGE branch prediction cluster in an 8-issue RVA23 RISC-V processor, from research through timing closure at 2.75 GHz. He holds +20 issued patents in processor design, architecture, and hardware virtualization. He is the author of Pacino and the uarchlabs methodology documented here.
Connect on LinkedIn.