Skip to content
FinObservatory

Valuation / Methodology

Valuation methodology

Three source families, their distinct vintages, and a register of everything in them that can mislead a reader who does not know it is there. Every number on /valuation and on the 94 industry pages comes from a query written against these files at build time.

The sources

TableCoverageSource
shiller_ie_data1,867 monthly rows, 1871-01 to 2026-07. S&P composite only.Robert J. Shiller, Irrational Exuberance data, Yale
damodaran_cost_of_capital, damodaran_betas, damodaran_valuation_multiples_{pe, pbv, ps, ev_ebitda}One vintage, 2026-01-05. 94 industries plus two aggregate rows, in two regions.Aswath Damodaran, NYU Stern
damodaran_erp_annual, damodaran_erp_monthly, damodaran_country_premium66 annual rows (19602025), 216 monthly rows (2008-09-012026-08-01), and 178 country rows in two source-defined blocks. Vintages 2026-01-08, 2026-08-01 and 2026-02-16.Aswath Damodaran, NYU Stern
kpss_patents3,419,552 patent rows, 1926 to 2024, 10,082 distinct permno.KPSS authors’ extended patent data

These two files do not describe the same moment. Shiller ends 2026-07; the Damodaran snapshot is dated 2026-01-05. No sentence on this layer combines a number from one with a number from the other.

The KPSS block is a third object again: a patent-level panel that the authors extended through 2024 and whose citation counts are updated through 2024. It is not joined to Shiller or Damodaran anywhere on the page.

CAPE

CAPE at month t is the real price of the index divided by the mean of the trailing 120 months of real earnings, both deflated to the last month’s price level. Recomputing it that way from real_price and real_earnings reproduces Shiller’s own cape column to within 0.325 index points at every one of the 1,746 months where 120 trailing earnings observations exist. The residual is not zero, so the published column is used, not the recomputation.

The denominator is a 120-month average by construction, which is the point of the measure and also its main limitation: CAPE responds to an earnings collapse over ten years, not over a quarter. cape is populated for 1,747 of the 1,867 rows, the missing ones being the first 120 months, before there is ten years of history to average. earnings_nominal is populated for 1,863 rows: it stops at 2026-03, so the last 4 months of the file carry a CAPE but no reported earnings. Today’s reading, 41.37 in 2026-07, is one of those months.

The forward-return test

real_10yr_ann_stock_return is Shiller’s own column: the annualised real return the index went on to deliver over the 120 months after each row. It is not computed here. Two null patterns bound the test:

  • cape is null for the first 120 months of the file, which need ten years of trailing earnings.
  • real_10yr_ann_stock_return is null for the last 120 months, which need ten years of future prices.

Both columns have the same non-null count, which makes them look aligned. They are not. 1,627 rows carry both, 1881-01 to 2016-07, and that is the only sample in which the test can run. The build asserts that the forward column ends exactly 120 months before the file and that cape begins exactly 120 months after it starts; if either stops being true, the build fails rather than the page quietly changing meaning.

Overlap

Two consecutive rows of the forward column describe windows sharing 119 of their 120 months. Its lag-1 autocorrelation is 0.993; at lag 12 it is 0.857. A t statistic computed on 1,627 such rows treats them as 1,627 independent draws, which is why it comes out at 23.7. That number is published on the index page next to its replacement so the size of the error is visible, not to be used.

The non-overlapping sample is drawn by taking the first testable month and every 120th month after it: 14 windows, 1881-01 to 2021-01, no two of which share a month. On those, t is 2.19 on 12 degrees of freedom. R-squared moves from 0.257 to 0.286, that is, hardly at all: overlap inflates the apparent precision of the estimate, not the estimate itself.

The regression reported is of the realised return on 1 / CAPE, the CAPE earnings yield, because a yield is the quantity that enters a return linearly. The index page reports CAPE and log CAPE alongside it so the choice can be inspected rather than trusted. Overlapping fit: a = 0.0069, b = 0.8594, residual SD 4.43% a year. Non-overlapping fit: a = 0.0092, b = 0.8108, residual SD 5.36% a year.

With 14 independent observations, the only available robustness check is to drop each in turn. Doing so moves R-squared between 0.106 and 0.376 and t between 1.14 and 2.57; 11 of the 14 refits fall below the two-sided 5 percent critical value of 2.201 at 11 degrees of freedom. That result is on the index page. It is the honest summary of what this test can support.

Damodaran equity and country risk premia

The annual chart selects implied_erp_fcfe, Damodaran’s “Implied ERP (FCFE)”. The file contains 66 year rows, but that column is blank in 1960, so the line contains 65 observations beginning 1961. The monthly chart selects implied_erp_t12m, his “ERP (T12m)”, from 216 month rows. No chart joins those definitions into one line or substitutes another ERP variant when a value changes.

These workbooks cannot use the industry files’ internal update-date cell. Their table vintages come from each HTTP Last-Modified header: 2026-01-08 for the annual history, 2026-08-01 for the monthly history and 2026-02-16 for the country cross-section. The build reads real openpyxl worksheets only, because the monthly workbook also contains a chartsheet, and coerces the text-formatted September 2024 date, index level, Treasury rate and ERP rather than dropping that row.

The country workbook is parsed as two tables. Rows in the rated block carry region, Moody’s rating and rating-based fields; rows in the frontier block carry a PRS score and do not inherit those rated-block columns. CDS-based figures are populated for 78 of 157 rated rows. The mature-market input of 4.23% equals the latest annual Implied ERP (FCFE) and latest monthly ERP (T12m). The US input of 4.46% equals the latest annual risk-adjusted ERP. The builder carries the source CRP rather than recomputing it: exactly 1 row, the United States, differs from direct subtraction by more than 1e-9 and none differs by more than 1e-4.

The Damodaran industry cross-section

94 industries, two regions, one vintage. The US tables cover US-listed companies; the Global tables cover the worldwide universe and are reported in US dollars, per the source file’s own header. Two further rows per region, Total Market and Total Market (without financials), are aggregates and are excluded from every cross-section here; the Total Market row is shown as a benchmark column on each industry page. The build asserts that the 94 industry firm counts sum exactly to the Total Market count in each region (5,994 US, 48,156 Global), so the aggregates cannot silently start double-counting.

Two identities, asserted at build time

  • Cost of equity is an exact affine function of beta within a region. US: 3.950% + 4.460% × beta. Global: 3.950% + 5.630% × beta. Maximum absolute residual 3.5e-17 and 5.6e-17 respectively. The build throws if the residual ever exceeds 1e-12. Read plainly: the intercept is the risk-free rate and the slope is the equity risk premium, applied identically to every industry, so the cross-section of cost of equity is the cross-section of beta rescaled.
  • Cost of capital equals its own weighted average. equity_to_capital × cost_of_equity + debt_to_capital × after_tax_cost_of_debt reproduces cost_of_capital. The build throws if the maximum absolute difference across the rows exceeds 1e-12.
  • Cost of debt is a bucket, not an industry estimate. It takes 4 distinct values across the 94 US industries and 4 across the 94 Global ones.

Traps in the source

  • Two PEs with the same name. current_pe, trailing_pe and forward_pe are averages across the firms in an industry. agg_mktcap_to_net_income_all_firms is total market capitalisation over total net income. Only the second is a market multiple. The industry pages and the index table lead with the aggregate.
  • Units are fractions, not percents. cost_of_capital, tax_rate, roe, roic, the margins and pct_money_losing_firms_trailing are stored as fractions (0.0696 means 6.96%). Betas and multiples are raw numbers. Every display here multiplies the fractions by 100 exactly once.
  • The two files are cross-checked against each other. damodaran_cost_of_capital.std_dev_in_stock and damodaran_betas.std_dev_of_equity are described in source data as the same quantity. They agree in all 192 shared rows (94 industries and two aggregates, in each of two regions), so the repaired files now line up exactly.
  • The multi-year beta columns are not shown. damodaran_betas carries yr_2022 through yr_2025 and average_beta_multiyear. The US spreadsheet titles the average column “Average (2022-2026)”; the Global spreadsheet titles the same column “Average (2020-24)”. The ingest normalised both to one column name, so a single label here would be wrong for one region. It is also not the arithmetic mean of the four year columns: across the 192 rows it equals that mean in 0 of them and the mean of those four plus the current beta in 0. All five columns are therefore omitted from the pages.
  • The industry names are Damodaran’s, spelling included. Names such as “Heathcare Information and Technology” and “Rubber& Tires” are reproduced as the source writes them. Correcting them here would break the join back to the source file.

Missing cells

All 35 data columns that these pages display were counted for nulls across the 188 industry-region rows of their tables. The 14 columns that have any null are listed in full; the other 21 are complete.

ColumnNullsIndustry (region)
agg_mktcap_to_net_income_all_firms10 of 188Broadcasting (Global); Drugs (Biotechnology) (Global); Real Estate (Development) (Global); Broadcasting (US); Chemical (Diversified) (US); Drugs (Biotechnology) (US); Electrical Equipment (US); Electronics (Consumer & Office) (US); Entertainment (US); Software (Internet) (US)
ev_ebit_all_firms9 of 188Bank (Money Center) (Global); Banks (Regional) (Global); Brokerage & Investment Banking (Global); Bank (Money Center) (US); Banks (Regional) (US); Brokerage & Investment Banking (US); Coal & Related Energy (US); Electronics (Consumer & Office) (US); Software (Internet) (US)
peg_ratio9 of 188Shipbuilding & Marine (Global); Chemical (Diversified) (US); Electronics (Consumer & Office) (US); Insurance (Life) (US); Paper/Forest Products (US); Real Estate (Development) (US); Real Estate (General/Diversified) (US); Reinsurance (US); Rubber& Tires (US)
ev_ebit_pos_ebitda_firms8 of 188Bank (Money Center) (Global); Banks (Regional) (Global); Brokerage & Investment Banking (Global); Bank (Money Center) (US); Banks (Regional) (US); Brokerage & Investment Banking (US); Coal & Related Energy (US); Electronics (Consumer & Office) (US)
roic8 of 188Bank (Money Center) (Global); Banks (Regional) (Global); Brokerage & Investment Banking (Global); Financial Svcs. (Non-bank & Insurance) (Global); Bank (Money Center) (US); Banks (Regional) (US); Brokerage & Investment Banking (US); Financial Svcs. (Non-bank & Insurance) (US)
ev_ebitda_all_firms7 of 188Bank (Money Center) (Global); Banks (Regional) (Global); Brokerage & Investment Banking (Global); Bank (Money Center) (US); Banks (Regional) (US); Brokerage & Investment Banking (US); Electronics (Consumer & Office) (US)
ev_ebitda_pos_ebitda_firms6 of 188Bank (Money Center) (Global); Banks (Regional) (Global); Brokerage & Investment Banking (Global); Bank (Money Center) (US); Banks (Regional) (US); Brokerage & Investment Banking (US)
agg_mktcap_to_net_income_profitable_firms3 of 188Chemical (Diversified) (US); Electronics (Consumer & Office) (US); Rubber& Tires (US)
expected_growth_5yr3 of 188Chemical (Diversified) (US); Real Estate (General/Diversified) (US); Reinsurance (US)
std_dev_operating_income_10yr3 of 188Bank (Money Center) (US); Brokerage & Investment Banking (US); Electronics (Consumer & Office) (US)
trailing_pe3 of 188Chemical (Diversified) (US); Electronics (Consumer & Office) (US); Rubber& Tires (US)
roe2 of 188Retail (Building Supply) (US); Tobacco (US)
current_pe1 of 188Electronics (Consumer & Office) (US)
price_to_book1 of 188Tobacco (US)

Nulls are rendered as “n/a” and never imputed, never zero-filled, and never dropped silently from a mean. Each industry page names the fields that are null for that industry.

The KPSS innovation block

kpss_patents is the authors’ patent-level panel: 3,419,552 patents from 1926 to 2024,matched to 10,082 distinct CRSP permno identifiers. The file is used only for aggregate views on /valuation: yearly totals, yearly value distribution, yearly count of distinct patenting firms, and citation buckets. No patent-level display and no firm naming are published from it.

  • Units are millions of dollars. The authors define xi_nominal as the value of innovation in millions of nominal dollars and xi_real as the same measure deflated to 1982 dollars using the CPI. The page labels the yearly sum in billions of 1982 dollars and the per-patent distribution in millions of 1982 dollars, which is only a display-scale change.
  • The yearly sum is shown beside the distribution because the tail is extreme. Over the full sample the median patent value is 3.545 million 1982 dollars, the 90th percentile is 27.4 and the 99th percentile is 132.2. The yearly table therefore keeps the sum beside the median, p90 and p99, rather than letting the total stand alone.
  • Citation buckets stop at 2019. The authors’ README says the forward-citation field is updated through 2024. A patent issued in 2023 or 2024 has had little or no time to accumulate citations, so the citation table excludes issue years after 2019. That gives every patent in the displayed cut at least five calendar years to be cited in the source file.
  • The identifier is the limit. permno is a CRSP security identifier, not a public company name. The estate carries no license-clean permno crosswalk, and the methodology forbids inventing one by any side route. The page is aggregate only for that reason, not because the dataset is thin.

What this layer will not do

  • Compare CAPE across countries. CAPE is a price index over a ten-year average of reported earnings; two markets with different index composition and different accounting rules produce numbers that cannot be ranked against each other. Only the S&P composite is shown.
  • Slice Damodaran through time. One vintage, 2026-01-05, and no other date column in the file.
  • Name firms from KPSS. The file carries CRSP permno and this estate carries no license-clean permno-to-name crosswalk.
  • Treat recent citations as settled. The citation table stops at 2019 because the source updates citations only through 2024.
  • Put a Shiller number and a Damodaran number in the same sentence. They are 2026-07 and 2026-01-05.
  • Claim that 14 observations settle the question. They do not, and the leave-one-out table on the index page shows how far they are from settling it.