Cockpit

THE FACTOR LAB HANDBOOK

The complete guide · v0.2.1

Glossary

Welcome — what this site is (and is not)

Every day, a computer reads the financial reports and price history of roughly 6,600 US stocks and ranks them on one leaderboard. Top of the leaderboard = the most evidence in the stock's favor. That's it. This handbook walks through exactly how that evidence is assembled, what every number and chip on screen means, and what the system is deliberately not doing.

Nothing here buys or sells anything, and nothing here is financial advice. It is a research shortlist with the evidence laid out — a machine for narrowing 6,600 stocks down to a shortlist worth your research time. Trades are never executed, not even in the paper portfolios (they are simulated, honestly, with real prices and real costs).

If a word has a dotted underline, it is a -level technical term — click it to see what it means without leaving the page. A complete, searchable index of every term lives in the glossary section.

What happens every day — the pipeline

Behind the page is an ordered chain of jobs, re-run automatically every day via . An orchestrator enforces the order and checks data integrity so stale or partial data can never silently corrupt your rankings.

⚙️The daily data pipeline

RUNS DAILY
① Parallel inputs — data sourcesParallel

SEC Company Facts

Ten years of fundamentals straight from — the ground truth behind , checks, and .

Yahoo Finance

Live prices, price history, analyst , and metadata for every ticker.

FRED (Fed macro)

Federal Reserve macro series that feed the used to trigger .

② Serial — the scoring chainSerial
1

Build fundamentals history

Turn raw filings into a clean 10-year F-score, Sloan , a real M-score, and net issuance. Missing values stay null; they are never silently treated as safe.

outputfundamentals_battery.json
2

Reverse engine

A first-pass safety & quality screen: classify the archetype (A–F), score survivability and data quality, and compute flags. It also writes candidates into an append-only nomination log.

outputsafety inputs · nominations
3

Factor Lab

The heart of the rankings. Every sub-metric is d and d within the stock's own group, averaged into the five s, ed, discounted by safety s, then cut into s — with hard es for disqualifying flags.

outputcomposite · band · rank
③ Parallel — two reads of the same listParallel
4

Valuation models (reverse DCF)

Solve a by for the growth the current price already assumes, then compare it with what the company has actually delivered. The result is the (DCF gap).

outputexpectations gap
5

RS2 LLM overlay

A local reads each company's actual filings and writes an independent verdict — , , — which can promote, demote, or a name after the quant bands are set.

outputLLM rank · verdict
6

Portfolio plan

Turn the Research Now list into a sized, capped allocation. -based , and , , and exit rules. Decision support only — nothing here executes trades.

outputsuggested plan (plan / plan2)
④ Multiple output destinationsOutputs

Rankings / leaderboard

The , , rank, and DCF gap you see on the dashboard, refreshed every run.

Portfolio plan

The (value core) and (hybrid) allocations with sizing, flags, and macro de-risk.

Track Record

Four portfolios daily with real , benchmarked against and — the honest meter.

AI Reports

On-demand analysis stored in Supabase and shown under /reports.

Forward logs → outcomes

Every signal and nomination is logged and later measured against real forward returns — no hindsight, no editing.

The whole chain re-runs daily via GitHub Actions.The IC drift report recalculates monthly.

On-demand AI analysis (Ask AI)

on-demand
1Ask AIYou click “Ask AI” on any stock card or detail view.
2/api/analysisqueues a pending job in Supabase.
3GitHub Actionstriggers the AI worker on a runner.
4DeepSeekreads the full RS2 prompt + the stock's financials.
5Supabasestores the verdict for polling.
6/reportsshows the finished analysis.

DeepSeek writes the analysis from the full RS2 prompt + financials; results appear under AI Reports (/reports).

Everything is : each logged signal uses only information that existed at that moment. That discipline is what makes the track record trustworthy.

Reading the leaderboard — the Rankings tab

The Rankings tab is a table of the whole universe, ranked by one number: the . Here is what every column means.

  • Rank — today's position on the leaderboard of ~6,600 scored US stocks. #1 has the strongest overall evidence right now.
  • Composite — the one number everything is ranked by (0–100). It blends the five factor scores, each measured against the stock's own comparison group, then applies safety s. Higher = more evidence in the stock's favor.
  • Factor mix — what is driving the score. A contribution bar shows the share of each factor: green = value, blue = quality, amber = momentum, violet = low volatility, pink = revisions. Longer segment = bigger contribution.
  • Band — what the rank means in practice (, , , ). A red chip means the stock is ed outright — the reason is written on the chip.
  • Market cap, the price of the whole company (share price × shares).
  • RS2 rank / stance / conviction / action — the independent AI read. See the RS2 section.
  • Δ pctl — how much the quant engine and the AI disagree, in percentile points. Big gaps are the interesting rows: one of them is wrong.
  • DCF gap — the : the growth the price requires vs the growth the company has actually delivered. See the DCF section.
Click any row for the per-stock detail: the full factor profile, the reverse-DCF read (the growth the price implies vs what the company has demonstrated), and the RS2 local-LLM research and verdict.

The five factors — the ingredients of a score

Each stock is graded on five traits that have historically predicted returns. Crucially, every grade is — a supermarket competes with supermarkets, not with software companies. Otherwise “high momentum” would just mean “is a tech stock.”

Value

Are you paying $1 for $2 of yearly cash earnings, or $2 for $1? Cheap beats expensive on average over time. Measured as the average of four yields — , , , and — each against the stock's current price. See the methodology.

Quality

Does the company make real money, consistently, with clean accounting? Combines , stability across several years, (preferring cash-backed earnings), the , and . A profitable business with honest books beats a story.

Momentum

Has the stock been winning over the past year? Winners tend to keep winning for a while. Built from the (the academic 12-month return skipping the last month) and . See for why the last month is skipped.

Low volatility

Does the price move calmly or wildly? Calm stocks have historically delivered more return per unit of pain. Measured as the negative of the standard deviation of monthly returns (at least 12 observations). The is also exported for sizing.

Revisions

Are the professional analysts who follow the company raising or cutting their forecasts? Direction of change matters. Built from the normalized slope and a structured score.

Why every ingredient counts equally

Think of judging a decathlon: you could try to guess which event matters most, but decades of research show those guesses backfire — the “perfect” weights found in past data almost never work on future data. So each of the five ingredients counts exactly the same (). Boring, humble, and it works better. The site still -measures each factor's predictive power every month as a diagnostic — it just never lets a short sample steer the engine.

A missing factor does not silently wreck a stock: if value, quality, or momentum is missing, the name is marked insufficient factors rather than scored on partial data. When a non-essential factor is missing, the remaining weights renormalize so the composite stays comparable.

Bands, vetoes & safety haircuts

The composite percentile is cut into four practical buckets, or s:

  • RESEARCH NOW — top 3%. Worth your research time today.
  • WATCHLIST — top 10%.
  • MONITOR — top 30%.
  • PASS — the rest.

Vetoes — hard disqualifiers

Some stocks are disqualified no matter how good the score looks — think of a house with beautiful photos that failed the structural inspection. A is applied before scoring, so a vetoed name gets no composite at all. Reasons include: the reverse engine's safety checks (reverse_engine_reject), both forensic alarms firing together ( + high ), from heavy share issuance, or (in the AI lens) a hard avoid/sell verdict. The reason is written on the red chip.

Safety haircuts

Between the raw score and the final rank, three multiplicative s are applied: survivability = 0.7 + 0.3 × (survivability/100); data quality = min(1, 0.8 + 0.04 × dq); and forensic = 0.85 if a single Beneish or accruals alarm fired (both firing is a veto, not a haircut). The result is re-ranked, so fragile or suspicious names drop without being thrown out.

The expectations gap — what the price silently promises

Every stock price silently makes a promise about future growth. The extracts that promise as a number — the the market is charging you for — by solving (via ) for the growth rate that makes a standard equal the current price.

The DCF gap column compares that promise with reality: minus (the last 5 years of revenue/FCF growth from SEC filings), in percentage points.

  • Green / negative — the price promises LESS than the company has proven. A potential bargain: you are being paid not to believe the growth story.
  • Amber / positive — the price needs an acceleration nobody has demonstrated yet. You have to believe a story.
The gap is also the used by the suggested plan: only names priced below their demonstrated growth have measurable edge, so high-ranked-but-expensive names are skipped with “no Kelly edge.”

The Lens — whose eyes you look through

At the top of the Rankings tab is one switch that decides whose ranking you see:

  • Quant — the deterministic factor engine. Pure math over financial statements and prices. No AI involved. This is the original, unchanged view.
  • RS2 LLM — the local AI analyst's own list, built by reading each company's actual filings and writing an independent verdict.
  • Compare — both side by side, biggest disagreements first. A disagreement percentile (Δ pctl) is computed per name; big gaps are where one engine is wrong.

The AI applies after the quant bands are set: it can promote, demote, or veto names, producing a parallel ranking. The quant baseline is never overwritten — the two lists are both shown so you can see where they disagree.

RS2 — the AI second opinion

RS2 is a that reads each company's actual SEC filings and writes an independent verdict — like getting a second doctor's opinion. For every name it produces:

  • — undervalued / fair / overvalued (a colored pill).
  • — how confident it is, 0–15.
  • — buy / hold / reduce / avoid… an opinion for research, never an order.
  • Its own DCF read, whose is anchored to the analyst consensus band and de-forwarded to — so the measures cheapness today, not a 12-month price target.

The AI Research Now gate

The AI's Research Now list keys off RS2's structured signals — margin of safety and entry timing — in two tiers: deep value (MoS ≥ 30%) earns Research Now at any conviction; moderate value (MoS ≥ 15%, or a genuine fresh buy) additionally needs conviction ≥ 9.5. Bearish calls are demoted out of Research Now; a hard avoid/sell is vetoed.

When a name leaves the quant Research Now list, RS2 writes an — a hold/trim/sell call for current holders, shown as an amber “LLM EXIT” chip.

Track Record — the honest meter

Instead of showing a flattering , the system s its own picks every single day with real prices and real , and the record is append-only — it can never be edited. If the machine is wrong, this page will say so, publicly and permanently. That's the point.

The portfolios

  • plan — the value core: -sized, ~50% cash.
  • plan2 — the hybrid: value core + , ~78% invested, holds the expensive leaders.
  • equal — equal-weighting every Research Now name (pure stock-picking test).
  • mine — your saved My Portfolio holdings, -measured like a fund.

How to read it

  • plan vs plan2 — if plan2 wins, paying up for quality leaders beat the value discipline this period; if plan wins, the discipline (and cash) paid off.
  • plan vs equal — the sizing machinery adds value if plan beats equal weighting.
  • equal vs IWM — the stock selection itself works if the picks beat the small-cap .
  • mine vs plan — your own deviations cost money: that's the .
  • “Sold too early” flags — exits that kept rising. A recurring pattern there means the exit rule needs work.

and appear only after enough days of live data; early on this page is deliberately boring. A what-if overlay lets you re-cost every trade at your own commission rate to see the drag of fees.

Portfolio — sizing & the suggested plan

The Portfolio tab has two very different halves. Read the labels carefully:

My Portfolio (top) — your actual holdings

Enter your ACTUAL holdings (saved only in this browser). Each is checked against the model: a quarter- suggested size, an over/under-weight verdict, and loud flags if a holding is ed or outside coverage. The “mine” ledger on Track Record uses these, unitized like a fund.

Suggested plan (below) — NOT your portfolio

A machine-built allocation from the Research Now list, with a Value core / Hybrid toggle:

  • Value core (plan) — quarter-Kelly sizing: = the expectations gap closing over ~3 years; risk = ; f = 0.25 × edge/risk², capped 5%. flags halve size; GPR 2–3 and insider selling shrink it; sector 25% / theme 30% caps; the rest stays (often ~50%).
  • Hybrid (plan2) — the same value core PLUS a that buys the top-ranked names REGARDLESS of valuation gap (capped ~35% of ), so it holds the expensive leaders the core refuses and deploys the idle cash (~78% invested). Sleeve rows are tinted pink.
Why two? The value core protects you in a bust (it won't overpay) but lags in a melt-up; the hybrid captures the leaders but rides them down harder (bigger s). Track Record shows how both actually perform.

Macro de-risk

are warning lights from data. If 2+ fire, every suggested size halves automatically () — shown as a loud amber banner on the tab.

Overlay chips & forensic flags

Overlays are context chips, never additive score. They exist to shrink positions, demand bigger margins of safety, or question your thesis:

  • GPR 0–3 tagged from the company's actual business profile (revenue geography, supply chains, regulation, sanctions). Never a buy/sell signal; at level 3 it shrinks position sizes and demands a bigger margin of safety.
  • ▲/▼ INSIDERS: insiders net-buying while short sellers retreat (▲, confirming) or insiders selling into elevated (▼, interrogate the thesis). Confirmation or warning only.

Forensic flags

The forensic battery — M-score, Sloan , net issuance, and the fundamentals battery — produces the flags shown on the plan rows. A single alarm is a 0.85 haircut; the pair firing together is a .

Themes — context, never a scoring factor

A is a market narrative a stock belongs to — AI, semiconductors, biotech, and so on. Theme membership and scores ride along for orientation and for crowding warnings (late-cycle theme crowding is a risk signal), but they never add to the composite.

This is deliberate: naive theme exposure has historically destroyed value — specialized theme ETFs average −3.1%/yr (Ben-David et al. 2023). The site won't let a hot narrative quietly inflate scores. In the plan, themes are bounded by a 30% theme cap so one hype story can't take over the book.

Methodology — for practitioners

Scoring pipeline — exact mechanics

Universe: every name scored by the reverse engine (~6,600 US listings). Per sub-metric: at the 1st/99th percentile within sector, then within sector. Factor z = mean of that factor's available sub-metrics. Composite z = weight-renormalized sum over available factors (missing factors drop out and remaining weights rescale; value, quality, and momentum are required — a name missing any of them is marked insufficient_factors rather than scored on partial data).

Composite z → cross-sectional (0–100) → three multiplicative s (survivability, data quality, forensic) → re-ranked → final percentile sets the : ≥97 research_now, ≥90 watchlist, ≥70 monitor, else pass.

Factor construction — sub-metrics and sources

Value = mean z of four yields, all computed from the latest fiscal year of SEC-filed fundamentals against current market cap: (FCF/mcap), ((NI + D&A − capex)/mcap), yield (operating income/), and (NI/mcap — broadest coverage, rescues filers with missing capex/D&A/op-income tags).

Quality = mean z of: (reverse-engine score), stability (−stdev across ≥4 fiscal years), negative (−accruals ratio), and (both from the forensic battery).

Momentum = mean z of the and . Monthly closes.

Low volatility = z of −σ(monthly returns), minimum 12 observations; the is exported per name and feeds sizing.

Revisions = mean of two 0–1 parts: normalized slope (clamp(slope, −1, 1)+1)/2, and analyst structured score/100 — scaled to 0–100 then re-centred to a z-like scale via (score−50)/25.

Weights

Equal 0.20 × 5 (scheme equal_weight_robust5). per factor is measured monthly but writes a drift diagnostic only — measured IC never steers the weights (DeMiguel, Garlappi & Uppal 2009: estimated weights rarely beat 1/N out of sample).

Reverse DCF — exact method

The valuation models solve by for the growth rate that makes a standard DCF equal the CURRENT price — the growth the market is charging you for. The (shown as “DCF gap”) = implied growth − demonstrated growth, where demonstrated = the last 5 years of revenue/FCF growth from SEC filings, in percentage points. The reverse engine layers archetype classification (A–F) and survivability/data-quality scoring on top, producing the safety inputs the Factor Lab consumes.

Veto rules (exact)

reverse_engine_reject = reverse-engine band ∈ {Excluded, Reject, Reject-tier}· forensic_pair = Beneish M-score elevated AND accruals high (single alarm = 0.85 haircut instead) · heavy_issuance = HEAVY_ISSUANCE flag, waived for archetypes E/F where issuance is the expected financing mode. In the AI lens, a hard avoid/sell verdict is also a veto (llm_reject).

Missing data is null, never silently safe

The old placeholder Z/M-scores are gone. If a metric is missing, it is null — and the scoring pipeline either marks the name insufficient or applies the data-quality haircut. A gap is never quietly treated as a pass.

How the system is validated

Two independent honesty loops keep the machine honest:

  • Forward-logged signals — every factor signal is logged at the moment it is made (, append-only) and later measured against what actually happened, including delisted names (no ).
  • Paper-traded portfolios — the Track Record page trades plan / plan2 / equal / mine daily with real prices and costs, benchmarked against and . Returns and are public and permanent.

Monthly, an IC drift report recalibrates the diagnostics. The point of all of this is that the site never gets to grade its own homework: the scoreboard is forward-looking, real, and uneditable.

Where the data comes from

  • — 10 years of as-filed fundamentals via (the ground truth for quality, forensics, and demonstrated growth).
  • — prices, analyst , and coverage.
  • — Fed macro series behind the .

The whole pipeline re-runs daily via ; the IC drift report recalculates monthly. Paper ledgers persist append-only with transaction costs in bps.

Frequently asked questions

Is this financial advice?

No. It is a research shortlist with the evidence laid out. Nothing here buys or sells anything, and nothing here is financial advice.

Does the site execute trades?

Never. Even the paper portfolios are simulated — but honestly, with real prices and real transaction costs.

Why is the composite ranked by five factors, not more?

Five robust, historically documented factors, equal-weighted on purpose. Adding more tuned factors invites , which loses to simple 1/N out of sample.

Why is the suggested plan often ~50% cash?

The value core only buys names with measurable — priced below demonstrated growth — and refuses to overpay. Cash is a feature: protection and dry powder.

What do the red chips mean?

A : automatic disqualification regardless of score. The reason is written on the chip.

Why do quant and RS2 disagree?

One is pure math over financial statements; the other reads filings with judgment. When they strongly disagree, one of them is wrong — those are the interesting rows. Use the Compare lens.

How can I trust the track record?

It is forward-logged (), append-only, cannot be edited, includes , and measures against real benchmarks including delisted names.

The glossary — every term in this handbook

A searchable index of all 106 defined terms. Click any term chip to open its definition — or click a term inline in any section above.

Core concepts16 terms
Bands & vetoes7 terms
Value factor6 terms
Quality factor9 terms
Momentum factor4 terms
Risk & volatility3 terms
Revisions factor5 terms
Valuation & DCF10 terms
RS2 / AI5 terms
Portfolio & sizing12 terms
Track record12 terms
Overlays & forensics8 terms
Data & pipeline9 terms

Disclaimer

This site is a research tool. It is not financial advice, not a trading bot, and not a crystal ball. Factors work on average over years, not on every stock every month. The system's own track record is honest and often humbling by design. Past performance — including paper-traded performance — does not guarantee future results. Do your own research.

Back to the Factor Lab Cockpit