Auditor Independence: Moving Toward Harmonization or Simplification?

INTRODUCTION

Auditor independence has been a priority for the Securities and Exchange Commission (“SEC”) under the leadership of both the Trump Administration and the Biden Administration. In 2020, former SEC Chair, Jay Clayton, pointed out that in the United States “auditor independence rules are far-reaching and restrictive,” which could have “unintended, negative consequences.”1Jay Clayton, Promoting an Effective Auditor Independence Framework, U.S. Secs. & Exch. Comm’n (Oct. 16, 2020), https://www.sec.gov/news/public-statement/clayton-promoting-effective-auditor-independence-framework-101620 [https://perma.cc/KZS5-WPMY]. Shortly thereafter, the SEC issued new regulation that lowered auditor independence requirements and brought the SEC’s independence rules closer to the rules set forth by the Public Company Accounting Oversight Board (“PCAOB”) and American Institute of Certified Public Accountants (“AICPA”), the other two regulatory entities responsible for auditor independence.2Press Release, U.S. Secs. & Exch. Comm’n, SEC Updates Auditor Independence Rules (Oct. 16, 2020), https://www.sec.gov/news/press-release/2020-261 [https://perma.cc/4Q63-BQEW]. Meanwhile, the current chair, Gary Gensler, has signaled that auditor independence remains a “perennial problem area,” indicating that a tightening of the auditor independence requirements is soon to be seen.3Gary Gensler, Chair, U.S. Secs. & Exch. Comm’n, Prepared Remarks at Center for Audit Quality “Sarbanes-Oxley at 20: The Work Ahead” (July 27, 2022), https://www.sec.
gov/news/speech/gensler-remarks-center-audit-quality-072722 [https://perma.cc/Y6QB-XJ4S].

While the future of auditor independence regulations remains up in the air, the problems associated with a lack of auditor independence continue. In 2019, the SEC alleged that PricewaterhouseCoopers LLP, one of the “Big Four” accounting firms, had violated auditor independence rules in connection with nineteen service engagements for fifteen publicly traded companies by providing prohibited non-audit services that could have impaired the firm’s objectivity.4Press Release, U.S. Secs. & Exch. Comm’n, SEC Charges PwC LLP with Violating Auditor Independence Rules and Engaging in Improper Professional Conduct (Sept. 23, 2019), https://www.sec.gov/news/press-release/2019-184 [https://perma.cc/MH4R-YR5Q]. Non-audit services are ancillary services, such as reviews of accounting software or tax advice, that do not assist in the goal of auditing, which is reviewing the financial statements for fraud or error. The SEC also alleged that another Big Four accounting firm, Ernst & Young LLP, had violated auditor independence standards, along with one of its partners and two of its former partners.5Press Release, U.S. Secs. & Exch. Comm’n, SEC Charges Ernst & Young, Three Audit Partners, and Former Public Company CAO with Audit Independence Misconduct (Aug. 2, 2021), https://www.sec.gov/news/press-release/2021-144 [https://perma.cc/HE8Q-MG2A]. The Chief Accounting Officer for Ernst & Young’s client on this engagement was also allegedly involved with this misconduct, indicating how far-reaching auditor independence violations can be.6Id.

Auditor independence is governed by a self-regulatory model, in which the SEC, in partnership with the PCAOB, the specialized non-profit corporation created by the Sarbanes-Oxley Act of 2002,7Further discussion of the Sarbanes-Oxley Act of 2002, including the sweeping reforms it encompassed, will be provided in Part I. provides oversight over the AICPA, which is a private industry professional organization charged with setting substantive auditor regulation. Despite the importance of auditor independence regulation in investor protection and the existence of this self-regulatory model, the regulatory framework in this area remains entangled.

While the SEC and PCAOB provide oversight over the AICPA, they also issue their own auditor independence regulation and have enforcement practices associated with auditors.8See infra Part III. Among the SEC, PCAOB, and AICPA, each standard-setter has rules that overlap with the others in the same subject-matter and some rules that defer to the rules set by the other organizations.9See infra Part III. This leads to a waterfall effect, in which all three entities have to change their regulations any time one of the other two does, in order to ensure that the rules are not in conflict.10See, e.g., Pub. Co. Acct. Oversight Bd., PCAOB Release No. 2020-003, Amendments to PCAOB Interim Independence Standards and PCAOB Rules to Align with Amendments to Rule 2-01 of Regulation S-X 3 (Nov. 19, 2020), https://pcaob-assets.azureedge.net/pcaob-dev/docs/default-source/rulemaking/docket-047/2020-003-independence-final-rule.pdf [https://perma.
cc/NW28-LB49].
This effort has been called “harmonization,” and has the goal of ensuring that all regulatory frameworks in this area are made consistent with each other, to provide more certainty to the accounting firms and other stakeholders involved in audits, including public company boards of directors.11See, e.g., Deloitte & Touche, Comment Letter on the Proposed Revision of the SEC’s Auditor Independence Requirements Regarding Scope of Services (Sept. 25, 2000), https://www.sec.gov/
rules/proposed/s71300/deloit1a.htm [https://perma.cc/WS68-UCEQ].
However, an alternative solution to harmonization may be “simplification,” in which instead of expending effort to harmonize the regulation of the three entities each time one of them makes a change, the existing self-regulatory model could be streamlined so that the AICPA is the primary or sole standard setter, with the SEC and PCAOB providing government oversight.

To explore whether simplification is a compelling alternative to harmonization, this paper turns to federal judicial decisions by conducting a novel case study of all auditor independence cases decided after the passage of the Sarbanes-Oxley Act of 2002. These decisions indicate that the courts largely use AICPA standards and case law requirements in assessing auditor independence, rather than SEC or PCAOB standards. This new finding suggests that simplification of the regulatory framework by relying solely on AICPA rulemaking is a viable solution, given that the federal courts already rely on AICPA rules.12See infra Parts IV and V.

This paper will proceed as follows. First, Part I provides a brief summary of the business of public company audits to preview how the structure of the audit industry gives rise to unique incentivizes and pressures that may impact auditor independence. Next, Part II includes an overview of the impact of the Sarbanes-Oxley Act of 2002 on the audit industry and highlights the debate over auditor independence, including the ways that various stakeholders have argued whether auditor independence is a worthy goal, or if auditors’ role as gatekeepers is unnecessary. Thereafter, Part III provides an overview of the self-regulatory model that governs auditor regulation, as well as regulations promulgated by the SEC, PCAOB, and AICPA in the area of auditor independence, including a summary of recent efforts to harmonize the standards set by each entity. The core of the paper’s contribution to the literature on auditor independence regulation is in Part IV, which a presents a novel case study that examines which body of regulation, between the SEC, PCAOB, and AICPA, is preferred by the federal courts when determining whether there have been auditor independence violations. This leads to the key finding that, in determining whether auditor independence violations have occurred, the federal courts rely almost exclusively on AICPA standards and case law requirements developed by the courts, rather than SEC or PCAOB standards. Accordingly, the paper then briefly examines in Part V whether this finding indicates that rule-making authority should be consolidated by giving the AICPA authority to set substantive auditor independence regulation, with the SEC and PCAOB providing oversight, given that there is already a self-regulatory model in place. Put differently, the paper concludes by considering whether the three regulatory frameworks should be simplified into that of the AICPA, the standard setter that is the most comprehensive, and the one that has been most acknowledged by the courts.

I.  THE BUSINESS OF PUBLIC COMPANY AUDITS

The role of auditors in public company financial reporting is to provide third-party reasonable assurance to investors that the financial statements of their client companies “are free of material misstatement, whether caused by error or fraud,” in the form of a formally issued audit opinion, which is appended to the clients’ public company SEC filings.13Auditing Standards § 1001.02 (Pub. Co. Acct. Oversight Bd. 2020). Audit opinions state whether the auditor believes that the financial statements included in the filings are free, in all material respects, from error or fraud.

Generally, auditors are viewed as “gatekeepers,” meaning individuals who are “reputational intermediaries who provide verification and certification services to investors.”14John C. Coffee Jr., Understanding Enron: “It’s About the Gatekeepers, Stupid,” 57 Bus. Law. 1403, 1408 (2002). As gatekeepers, auditors have incentives to signal to outsiders that they are credible because their reputation of trustworthiness is what allows them to attract future clients and remain in business by giving credibility to their audit opinions.15Ronald J. Gilson & Reinier H. Kraakman, The Mechanisms of Market Efficiency, 70 Va. L. Rev. 549, 607 n.166 (1984). The reputation of auditors is also partially what enables them to provide verification services because investors recognize that auditors have fewer  incentives than their clients’ management teams to mislead investors because the management teams are corporate insiders whereas auditors are objective third parties.16Coffee, supra note 14, at 1406. A traditional understanding of the audit profession puts forth the proposition that auditors are trustworthy because they are not incentivized to risk their reputation by assisting one client with fraud, which could lose them many additional clients and destroy their reputational capital.17Id.

Additionally, auditors have built up years of expertise that demonstrates to third parties that their services and opinions can be trusted.18Id. at 1408. The utility of auditors extends from the fact that they have specialized technical expertise in the area of accounting. Members of accounting firms that conduct audits are generally expected to hold state licenses as Certified Public Accountants, which give them both reputational capital and technical expertise.19See Auditing Standards § 1010 (Pub. Co. Acct. Oversight Bd. 2020). Additionally, these accounting firms are typically well staffed with local and global teams to address the audit needs of multinational companies that are listed on exchanges in the United States, despite the complexity and geographical scope of the audit procedures that may be required.20Pub. Oversight Bd., Panel on Audit Effectiveness, Report and Recommendations 157 (2000), https://egrove.olemiss.edu/cgi/viewcontent.cgi?article=1351&context=aicpa_assoc [https://
perma.cc/TH2E-YLG2] [hereinafter Panel on Audit Effectiveness].

It is widely known that auditing is part of an oligopoly, in which the “Big Four” accounting firms—PricewaterhouseCoopers LLP (“PwC”), Ernst & Young Global Limited (“Ernst & Young”), Deloitte Touche Tohmatsu Limited (“Deloitte”), and KPMG International Limited (“KPMG”)—dominate the market. The Big Four were responsible for the audits of 88% of SEC large accelerated filers—which are companies with a public float of larger than $700 million that are required by the SEC to submit securities filings on a shorter timeline than other filers—and 44.7% of all public companies in 2022.21Nicole Hallas, Who Audits Public Companies – 2022 Edition, Audit Analytics (June 28, 2022), https://blog.auditanalytics.com/who-audits-public-companies-2022-edition [https://perma.cc/
4QSY-XMJ7; 17 C.F.R. §  240.12b-2(2) (2022).
There are several smaller players in the market as well, including RSM US LLP, BDO USA LLP, and Grant Thornton LLP; however, the largest of these entities produces less than a third of the revenue produced by the smallest Big Four accounting firm, leaving the market substantially dominated by the Big Four accounting firms, which have offices globally.22The 2021 Top 100 Firms, Acct. Today, https://www.accountingtoday.com/the-2021-top-100-firms-data [https://perma.cc/J5WV-7QVV].

Auditors for public companies are required to examine their clients’ financial statements and notes to the financial statements, which provide supplemental information. Auditors use a variety of mechanisms, including inspecting records, confirming balances with third parties, checking compliance with internal policies, and completing detailed tests of transactions to ensure the financial statements are free from material error or fraud.23See, e.g., Codification of Acct. Standards & Procs., Statement on Auditing Standards No. 110, § 318.79 (Am. Inst. of Certified Pub. Accts. 2006). Auditors also often test the robustness of management’s internal controls over financial reporting. Internal controls over financial reporting are policies and procedures surrounding accounting and reporting that are designed to limit the risk of fraud and error in the financial statements. All of this testing is done with an expectation of independence—that the auditors are not personally invested in the entity they are auditing, do not have conflicts of interests, and are reviewing the information with an air of “professional skepticism.”24See, e.g., Auditing Standards § 1015.07–09 (Pub. Co. Acct. Oversight Bd. 2020).

II.  THE DEBATE OVER AUDITOR INDEPENDENCE

When discussing auditor independence, much emphasis has been placed on the fact that each of the large public accounting firms have three lines of business: audit, tax, and consulting services. In the wake of the Enron scandal, debate over whether these three lines of business lead to inherent conflicts within public accounting firms that obstruct independence. Enron was a publicly traded energy company based in Texas that went bankrupt in 2001, partially as a result of a major decline in stock price after accounting irregularities were discovered at the company.25William W. Bratton, Enron and the Dark Side of Shareholder Value, 76 Tul. L. Rev. 1275, 1276, 1305–09 (2002). At the time, Enron was the seventh-largest company in the United States based on market capitalization, and its bankruptcy was the largest in United States history.26Id. at 1276. The accounting irregularities alleged were largely technical in nature, involving the use of (1) mark-to-market accounting, a practice that can lead to the overstatement of the value of assets by recording them at their current market value rather than their historical cost; (2) improper recognition of liabilities within “special purpose entities” that were owned by the company, which reduced the liabilities that were attributed to Enron; as well as (3) alleged improper payments to company officers, all of which were not in accordance with Generally Accepted Accounting Principles (“GAAP”) that are set by the AICPA and required to be followed by public companies.27Id. at 1282, 1305–09, 1348. Arthur Andersen LLP (“Arthur Andersen”), the public accounting firm responsible for Enron’s audit, did not, however, note any of these issues with GAAP in the audit opinions it issued about the accuracy of Enron’s financial statements, which led to surprise when the company collapsed.28See, e.g., Enron Corp., Annual Report (Form 10-K) (Mar. 30, 2001).

In the wake of the Enron scandal, many stakeholders argued that one of the main contributing factors to Enron’s collapse was that the company’s auditor, Arthur Andersen, was receiving more revenue from its consulting engagement with Enron than it was from its audit engagement, to a factor of 1.08 times.29Jonathan D. Glater, Enron’s Many Strands: Accounting; 4 Audit Firms Are Set to Alter Some Practices, N.Y. Times (Feb. 1, 2002), https://www.nytimes.com/2002/02/01/business/enron-s-many-strands-accounting-4-audit-firms-are-set-to-alter-some-practices.html [https://perma.cc/WHJ3-4QC9]. Arthur Andersen received approximately $27 million in annual non-audit fees and $25 million in annual audit fees from Enron in 2000 and expected to grow the overall fees to approximately $100 million annually, which some alleged was why Arthur Andersen did not disclose Enron’s lack of compliance with GAAP.30Id.; Thaddeus Herrick & Alexei Barrionuevo, Were Enron, Anderson Too Close to Allow Auditor to Do Its Job?, Wall St. J. (Jan. 21, 2002, 12:01 AM), https://www.wsj.com/
articles/SB1011565452932132000 [https://perma.cc/8FF9-H6CV].
Arthur Andersen was exonerated of any liability; however, the court did not reach the issue of whether the firm was conflicted and what effect that may have had on the audit.31Arthur Andersen LLP v. United States, 544 U.S. 696, 708 (2005).

Enron was used as an example of how pressure on accounting firms to maximize revenue from consulting engagements prevents accounting firms from robustly auditing financial statements. Primarily, stakeholders worried that accounting firms would fail to report material misstatements or fraud in financial statements in order to preserve relationships with management of the companies they were auditing or to sell them consulting services; in extreme cases such as Enron’s, the concern was that this failure to report could collapse the company entirely and thus completely eliminate a revenue-generating client.32Herrick & Barrionuevo, supra note 30.

The furor and fallout over Enron and implications on the weaknesses of accounting firms was used as a major justification for Congress’s passage of the Sarbanes-Oxley Act of 2002, given that investors lost billions upon Enron’s collapse.33Jeff Lubitz, 20 Years Later: Why the Enron Scandal Still Matters to Investors, Inst. S’holder Servs. Insights (Oct. 20, 2021), https://insights.issgovernance.com/posts/20-years-later-why-the-enron-scandal-still-matters-to-investors [https://perma.cc/A2DS-L7KR]. Specifically, the social costs of the Enron collapse were high because pension funds that supported teachers, firefighters, and government employees were endangered from the losses, leading to outcries for reform.34Steven Greenhouse, Enron’s Many Strands: Retirement Money; Public Funds Say Losses Top $1.5 Billion, N.Y. Times (Jan. 29, 2002), https://www.nytimes.com/2002/01/29/business/enron-s-many-strands-retirement-money-public-funds-say-losses-top-1.5-billion.html [https://perma.cc/H728-T2P8]; Legislative History of Title VIII of H.R. 2673, 148 Cong. Rec. S7419–20 (daily ed. July 26, 2002). William Donaldson, the SEC Chair during the time of the passage of Sarbanes-Oxley, testified before Congress that Enron was a major event leading to the reforms that were implemented in Sarbanes-Oxley.35Implementation of the Sarbanes–Oxley Act of 2002 Hearing Before the S. Comm. on Banking, Hous. and Urban Affairs, 108th Cong. 33–47 (2003) (statement of William H. Donaldson, Chair, Securities & Exchange Commission).

Sarbanes-Oxley included sweeping regulatory changes designed, in part, to “restor[e] public confidence” in the accounting profession, emphasizing for the first time in Congressional legislation the importance of auditor independence.36Id. The Act created the PCAOB to inspect accounting firms and set auditing regulations, departing from the prior regulatory structure in which the accounting profession was almost entirely self-regulated by the AICPA.37Id. Sarbanes-Oxley also created a slew of additional auditor independence rules, such as requiring audit committees to pre-approve all audit and non-audit services provided by an auditor, reducing the consulting services that could be provided by an auditor, requiring that certain auditors rotate off an audit engagement on a regular schedule, requiring that independence concerns be raised to the audit committee when auditors identified them, and requiring disclosure to investors of non-audit and audit services provided by accounting firms.38Id. The stated purpose of these reforms was not only to “restor[e] public confidence in the independence and performance of auditors of public companies’ financial statements,” but also to “enhance the integrity of the audit process and the reliability of audit reports.”39Id.

While Sarbanes-Oxley seemed to significantly reduce the opportunity for accounting firms to provide both audit services and consulting services to their clients, there exists a gray area in which accounting firms are able to provide certain “non-audit services” to their audit clients.40Office of the Chief Accountant: Application of the Commission’s Rules on Auditor Independence, U.S. Secs. & Exch. Comm’n, https://www.sec.gov/info/accountants/
ocafaqaudind080607#nonaudit [https://perma.cc/QC62-KZKF].
Accounting firms have found that these non-audit services can be a profitable substitute for revenue lost from consulting services, and according to a study that reviewed audit revenue as compared to non-audit revenue disclosures in annual proxy statements—which require public companies to disclose certain matters prior to their annual shareholder meetings—non-audit revenue represented 18% of the fees paid to auditors by public companies in 2021.41Nicole Hallas, Twenty Year Review of Audit & Non-Audit Fee Trends Report, Audit Analytics (Oct. 11, 2022), https://blog.auditanalytics.com/twenty-year-review-of-audit-non-audit-fee-trends-report [https://perma.cc/LT4F-32KW]. These non-audit services allow public company auditors to provide a number of ancillary services to their audit clients, such as, (1) audits over accounting and information systems, including payroll software and human resources software, (2) “carve-out” audits that evaluate portions of the business that are set to be divested, (3) tax services, (4) statutory audits, which are audits that comply with local governmental audit requirements, especially in foreign countries, and (5) employee benefit plan audits, typically meaning audits over pension plans. All of these services can be provided without impairing independence under the SEC, PCAOB or AICPA rules, and several of them are repeat, annual services, making them especially lucrative and enticing to accounting firms.42See Our Point of View on Non-Audit Services Restrictions, PricewaterhouseCoopers (2016), https://www.pwc.com/gx/en/about/assets/gra-non-audit-services-our-point-of-view.pdf [https://
perma.cc/PR9L-X8W2].

The oligopoly presented by the Big Four accounting firms also puts significant cost pressure on accounting firms to lower audit costs to attract and retain clients, while still maximizing revenue and market share.43Martin Gelter & Aurelio Gurrea-Martinez, Addressing the Auditor Independence Puzzle: Regulatory Models and Proposal for Reform, 53 Vand. J. Transnat’l L. 787, 798–99 (2020). This balance is primarily struck by the accounting firms in two ways.

First, the accounting firms can look to drive efficiencies in the audit process to lower the cost of annual audits to clients.44Id. at 803. This is often accomplished by retaining clients and using institutional knowledge gained over years of repeat audits to streamline audit processes and thus reduce hours and audit costs.45Id. at 808–09. However, this has the added effect of lowering auditor independence as auditors become entrenched in relationships with their clients over a number of years.46Richard L. Kaplan, The Mother of All Conflicts: Auditors and Their Clients, 29 J. Corp. L. 363, 367 (2004). For this reason, mandatory rotation of audit firms has been adopted in a number of foreign jurisdictions, including in Europe.47EU Audit Reform – Mandatory Firm Rotation, PricewaterhouseCoopers (2015), https://www.pwc.com/gx/en/audit-services/publications/assets/pwc-fact-sheet-1-summary-of-eu-audit-reform-requirements-relating-to-mfr-feb-2015.pdf [https://perma.cc/PY7R-9Q4D]. Some have also argued that mandatory audit firm rotations should be adopted in the United States, in addition to the existing United States requirement of rotation of the individuals who lead each audit project, known as audit engagement partners and who are usually equity partners in the accounting firms.4817 C.F.R. § 210.2-01(f)(7)(ii) (2023); see, e.g., Clive S. Lennox, Xi Wu & Tianyu Zhang, Does Mandatory Rotation of Audit Partners Improve Audit Quality?, 89 Acct. Rev. 1775, 1801 (2014).

Alternatively, accounting firms can look to increase revenue from ancillary, non-audit services.49Sean M. O’Connor, Strengthening Auditor Independence: Reestablishing Audits as Control and Premium Signaling Mechanisms, 81 Wash. L. Rev. 525, 559 (2006). This method of increasing revenue by cross-selling non-audit services along with public company audits is seen as a positive to many because efficiencies are driven by the institutional knowledge that auditors already have as a result of their financial statement testing, lowering the costs for both the audit and the non-audit services. For example, an accounting firm that is performing both a public-company financial statement audit and a report on internal controls for the same client can use some of the testing done for the financial statement audit to satisfy testing requirements for internal controls, thus lowering the costs for the combined services.50Auditing Standard § 2315.44 (Pub. Co. Acct. Oversight Bd. 2020).

Many argue that this bundling is a positive development, and given that audit services are fungible, lower transaction costs for companies in the form of reduced fees to accounting firms increase shareholder value.51O’Connor, supra note 49, at 542. Additionally, tracking independence over each and every ancillary service would increase transaction costs without providing shareholder value, given that some of these non-audit services can be provided without impairing auditor independence, and especially because some argue that accounting firms are already conflicted by virtue of the fact that they are paid for audit services provided.52Coffee, supra note 14, at 1411. Further, proponents of bundling argue that auditors are still incentivized to provide quality audits because despite the pressure to increase non-audit service revenue, auditors would not want to risk their reputational capital, which allows them to stay in business and attract other clients.53Gilson & Kraakman, supra note 15, at 607.

However, others argue that fee dependence on non-audit services poses the same problems as those that existed prior to Sarbanes-Oxley, in which instead of sacrificing audit quality to retain consulting revenue, firms are now sacrificing audit quality to retain other non-audit service revenue.54Paul Munter, The Importance of High Quality Independent Audits and Effective Audit Committee Oversight to High Quality Financial Reporting to Investors, U.S. Secs. & Exch. Comm’n (Oct. 26, 2021), https://www.sec.gov/news/statement/munter-audit-2021-10-26 [https://perma.cc/4P5E-UBX6]. This leads to concerns that auditor independence will be impaired in a way that increases costs to investors because material misstatements are less likely to be caught, and this outweighs the costs saved from efficiencies by the same accounting firm providing both audit and non-audit services.55Id.

While this paper reserves judgment on these questions of how non-audit services effect auditor independence, the primary view of auditor regulation since the accountings scandals of the early-2000s and Sarbanes-Oxley has been that auditor independence is of paramount importance to the audit industry in order to protect investors from fraud or inadvertent material errors from management.56Gensler, supra note 3. In the following Part, this paper will explore the existing regulation surrounding auditor independence, including identifying the many agencies and self-regulatory organizations that promulgate such regulations, and untangle the many requirements for independence as they have appeared over time.

III.  SELF-REGULATION AND THE TANGLED REGULATORY FRAMEWORK OF AUDITOR INDEPENDENCE

Modern third party audits have existed since approximately the 1920s and have varied widely in their form and requirements.57O’Connor, supra note 49, at 526–28. However, it was not until the Securities Act of 1933 that companies were required to have their financial statements certified by an independent accountant before they could register on the public markets.58Id. at 530, 535. This requirement was expanded with the passage of the Securities Exchange Act of 1934, which mandated annual and quarterly financial reporting for public companies, as well as reporting on material events, all of which were required to be certified by “independent public accountants.”59Panel on Audit Effectiveness, supra note 20, at 109. The Federal Trade Commission and SEC subsequently passed rulemaking that defined “independent public accountants” as those who had neither served as officers or directors of the company to be audited, nor had a “substantial financial interest” in the company, which was defined as more than one percent of the auditor’s net worth.60Id.

While the SEC retained statutory authority over auditor independence regulations, eventually the AICPA formed its own auditor independence requirements and in 1977 created the Public Oversight Board (“POB”) to serve as a self-regulatory organization over the accounting profession.61About the POB, Pub. Oversight Bd., https://www.publicoversightboard.org [https://
perma.cc/WC69-ALWN].
The AICPA is independently funded by membership dues from members, which include individual accountants who are certified public accountants.62Bylaws and Implementing Resolutions of Council § 2.3.1, Am. Inst. of Certified Pub. Accts. (May 19, 2022), https://www.aicpa.org/resources/download/aicpa-bylaws-and-implementing-resolutions-of-council [https://perma.cc/A8QH-CTUU]. The AICPA is governed by a council consisting of AICPA members elected by their fellow members in each state, representatives of state CPAs, and members-at-large of the AICPA, among others.63Id. § 3.3.1. However, in the wake of the accounting scandals of the early 2000s, the SEC issued additional independence rules to move away from complete industry self-regulation by the AICPA and POB without public oversight, and these rules were then amended to coincide with the sweeping reforms passed by Congress in the Sarbanes-Oxley Act.64Press Release, U.S. Secs. & Exch. Comm’n, SEC Adopts Rules on Provisions of Sarbanes-Oxley Act (Jan. 15, 2003), https://www.sec.gov/news/press/2003-6.htm [https://perma.cc/QTX5-T8CB]. Sarbanes-Oxley required the formation of the PCAOB—a nonprofit corporation under the oversight of the SEC—but still allowed the AICPA to issue auditor independence standards.65See 15 U.S.C. § 7211(a). The membership of the PCAOB, which consists of a five person Board, is determined by appointment from the Chair of SEC and a vote of the five SEC commissioners.66The Board, Pub. Co. Acct. Oversight Bd., https://pcaobus.org/about/the-board [https://perma.cc/KP8Q-VJ2G]. The PCAOB is funded by mandatory fees from public companies who are subject to audit requirements under the Exchange Act as well as fees from brokers and dealers subject to SEC regulation.67Accounting Support Fee, Pub. Co. Acct. Oversight Bd., https://pcaobus.org
/about/accounting-support-fee#:~:text=The%20largest%20source%20of%20funding,audited%20by%20
PCAOB%2Dregistered%20firms [https://perma.cc/34FM-75AH].
As a result, since 2003, when Sarbanes-Oxley took effect, the accounting profession has been governed by three primary sets of auditor independence standards, those set by the SEC, PCAOB, and the AICPA.68See O’Connor, supra note 49, at 559. A summary of the interaction of the regulations set by the AICPA, SEC, and PCAOB is provided below and will be discussed further in the following sections:

Figure 1.  

A.  Self-Regulatory Models in United States Financial Regulation

The diffuse structure that contains SEC, PCAOB, and AICPA involvement in auditor independence regulation is not unheard of within United States financial services regulation. The SEC is involved in several self-regulatory models in which the SEC provides oversight to an industry organization that sets standards, and, in some cases, enforces those standards. For example, the Financial Industry Regulatory Authority (“FINRA”) sets regulation for registered broker-dealer firms and registered brokers69Broker-dealer firms and brokers are organizations and individuals, respectively, who purchase and sell securities on behalf of their customers. and enforces those rules.70What We Do, Fin. Indus. Regul. Auth., https://www.finra.org/about/what-we-do [https://perma.cc/VX74-AZ7J]. The Exchange Act gives the SEC the authority to approve any rule changes made by FINRA7115 U.S.C. § 78s (b)(2). as well as revoke the authority of FINRA or issue sanctions against FINRA.72See id. § 78s (g)–(h).

The SEC is also part of self-regulatory frameworks with national securities exchanges.73National securities exchanges are exchanges for securities that have registered with the SEC under Section 6 of the Securities and Exchange Act of 1934. The self-regulatory framework between the national securities exchanges and the SEC works much in the same way as SEC’s oversight of FINRA, in which the SEC provides oversight over and approval of the rulemaking of each securities exchange and issues sanctions for violations of the Securities Act of 1933, the Investment Advisers Act of 1940, and the Investment Company Act of 1940.7415 U.S.C. § 78s(h)(2)(A).

The SEC also has a self-regulatory model with respect to credit rating agencies.75Under the Credit Rating Agency Reform Act of 2006, credit rating agencies are organizations that are engaged in the business of issuing an assessment of the creditworthiness of an obligor, using qualitative or quantitative models, and receiving fees for those services. Pursuant to Section 15E of the Securities and Exchange Act of 1934, each credit rating agency was required to set internal controls and policies to ensure accurate credit ratings,7615 U.S.C. § 78o–7(c)(3)(A). and the SEC was required to review such policies to ensure they were robust.77Id. § 78o–7(c)(1), (d)(1). Additionally, prior to the passage of the Dodd-Frank Wall Street Reform and Consumer Protection Act of 2010 (“Dodd-Frank Act”), the SEC relied on credit agencies to provide a measure of credit-worthiness in SEC rules.78See U.S. Secs. & Exch. Comm’n, Report on the Review of Reliance on Credit Ratings 1–2 (2011), https://www.sec.gov/files/939astudy.pdf [https://perma.cc/A5RR-BF5P]. However, Section 939A of the Dodd-Frank Act required that the SEC remove such references to the credit agencies within its rules and instead create its own independent standards.79Id.; Dodd-Frank Act, Pub. L. No. 111-203, § 939A, 124 Stat. 1887 (2010).

Although the examples above illustrate that the self-regulatory model has been widely used, the model has also been subject to debate over its merits.80See, e.g., A Review of Self-Regulatory Organizations in the Securities Markets Hearing Before the S. Comm. on Banking, Housing, and Urban Affairs, 109th Cong. (2006), https://www.
govinfo.gov/content/pkg/CHRG-109shrg39621/html/CHRG-109shrg39621.htm [https://perma.cc/93P2-45PD].
Supporters of the model argue it is beneficial because the knowledge of industry participants enhances the usefulness of rulemaking,81Andrew F. Tuch, The Self-Regulation of Investment Bankers, 83 Geo. Wash. L. Rev. 101, 112–13 (2014). places the costs of enforcement on industry,82See, e.g., Accounting Support Fee, Pub. Co. Acct. Oversight Bd., https://pcaobus.org
/about/accounting-support-fee#:~:text=The%20largest%20source%20of%20funding,audited%20by%20
PCAOB%2Dregistered%20firms [https://perma.cc/34FM-75AH].
and is more adaptive than the state-driven model because industry entities have the experience and ability to focus on specialized regulations in a way that public entities with vast oversight responsibilities may not.83Tuch, supra note 81, at 112. However, others believe that the self-regulatory model is not as beneficial as the European state-driven model, in which public entities are the singular or primary regulators with little or no input from industry organizations, because the self-regulatory model can result in conflicts of interest in terms of funding and regulatory capture, given that individual standard setters may be members of the group facing regulation and may seek to avoid restrictions.84See Saule T. Omarova, Bankers, Bureaucrats, and Guardians: Toward Tripartism in Financial Services Regulation, 37 J. Corp. L. 621, 628–29 (2012). In response, proponents of self-regulation may argue that the oversight of the public entities is sufficient to mitigate these conflicts and harness the benefits of a specialized industry standard-setter working hand-in-hand with a public oversight agency, especially when safeguards such as individual term limits for regulators, or limitations on the ability of individual regulators to participate in a “revolving door” in which they rotate between industry and regulatory entity, are put in place.85See Revolving Door Rules, Fin. Indus. Regul. Auth., https://www.finra.org/
careers/alumni/revolving-door-rules [https://perma.cc/Y2QH-NQL7].

While the self-regulatory model is not unique to auditor regulation, what is somewhat distinctive about the auditor regulation self-regulatory model is that the SEC and PCAOB not only provide oversight over the AICPA, which is the industry organization run by accountants, but the SEC and PCAOB also have concurrent jurisdiction to set independence standards and have set their own standards for auditor independence in addition to those of the AICPA.86Pub. Co. Acct. Oversight Bd., Release No. 2020-003, Docket No. 047, Amendments to PCAOB Interim Independence Standards and PCAOB Rules to Align with Amendments to Rule 2-01 of Regulation S-X 2–3 (Nov. 19, 2020), https://pcaob-assets.azureedge.net/pcaob-dev/docs/
default-source/rulemaking/docket-047/2020-003-independence-final-rule.pdf?sfvrsn=43d58c7e_6 [https
://perma.cc/GGD3-LTLC].
While there are additional organizations that set auditor independence standards, including state boards of accountancy, state CPA societies, federal and state agencies, and the International Ethics Standards Board, the standards set by these entities are heterogeneous and are generally superseded by the SEC, PCAOB, and AICPA rules.87Am. Inst. of Certified Pub. Accts., Plain English Guide to Independence 3 (2021), https://us.aicpa.org/content/dam/aicpa/interestareas/professionalethics/resources/tools/downloadabledocuments/plain-english-guide.pdf [https://perma.cc/N4AS-E3NG]. See generally O’Connor, supra note 49. Therefore, the regulation promulgated by these entities will not be examined in this paper. In order to assess the current landscape of auditor regulation, this paper will examine the regulations passed by the SEC, PCAOB, and AICPA in Sections III.B–D.

B.  SEC Independence Rules

Under the current SEC independence rules:

The [SEC] will not recognize an accountant as independent . . . if the accountant is not, or a reasonable investor with knowledge of all relevant facts and circumstances would conclude that the accountant is not, capable of exercising objective and impartial judgment on all issues encompassed within the accountant’s engagement. In determining whether an accountant is independent, the [SEC] will consider all relevant circumstances, including all relationships between the accountant and the audit client.8817 CFR § 210.2-01(b).

The SEC independence rules then set forth a non-exhaustive list of instances in which an auditor would not be independent, primarily consisting of eight categories: (1) lack of financial independence, (2) client investment in the accounting firm, (3) employment by client, (4) non-audit services, (5) contingent fees,89Contingent fees are:

[A]ny fee established for the sale of a product or the performance of any service pursuant to an arrangement in which no fee will be charged unless a specified finding or result is attained, or in which the amount of the fee is otherwise dependent upon the finding or result of such product or service.

Id. § 210.2-01(f)(10). In the context of the audit, contingent fees are generally understood to be fees made conditional on the finding of an “unqualified” audit opinion, in which the accounting firm certifies that the financial statements are reasonably and fairly presented.
(6) improper partner rotation, (7) lack of audit committee approval,9015 U.S.C. § 78c(a)(58)(A). Audit committees are committees established by the Boards of Directors of public companies which are charged with the oversight of financial reporting and audits. and (8) improper compensation.9117 CFR § 210.2-01(c). Perhaps the most important of these rules are those defining financial independence because without financial independence, an auditor may have conflicts of interest that prevent them from conducting a robust audit.92Auditor Independence Matters, U.S. Secs. & Exch. Comm’n, https://www.sec.gov/page/oca-auditor-independence-matters [https://perma.cc/8F2S-B7TB] (“Ensuring auditor independence is as important as ensuring that revenues and expenses are properly reported and classified.”). Accordingly, a summary of these rules is given below, along with a brief summary of the independence restrictions posed by the remainder of the SEC independence rules.

1.  Financial Interests

SEC Independence Rule 2-01 states that independence is impaired when (1) the accountant has a direct financial interest or material indirect financial interest in their client;9317 C.F.R. § 210.2-01(c)(1). (2) the accounting firm, a covered person in the firm, or any of the covered person’s immediate family members have a direct investment in the client, in which “covered person” includes individual accountants within a firm that provide services to a client;94Id. § 210.2-01(c)(1)(i)(A). (3) any partner or employee in the firm, including their close family, has more than 5% beneficial ownership of the client’s securities or controls the client;95Id. § 210.2-01(c)(1)(i)(B). or (4) the accounting firm, a covered person in the firm, or any of the covered person’s immediate family members have loans, savings accounts, checking accounts, broker-dealer accounts, insurance products, futures commission merchant accounts, consumer loans, or financial interests in investment companies that own the client.96Id. § 210.2-01(c)(1)(ii)(A)–(E).

In addition to the requirements of financial independence listed above, the SEC independence rules also prohibit the audit client from investing in the accounting firm or underwriting an accounting firm’s securities.97Id. § 210.2-01(c)(1)(iv). Moreover, the SEC independence rules place limits on accounting firm employees from being employed at the client both during and after the audit engagement.98Id. § 210.2-01(c)(2).

2.  Audit Conduct

The SEC independence rules also mandate certain conduct during the audit. For example, auditors are not allowed to provide non-audit services that could impair their independence, such as bookkeeping services, financial information systems design and implementation, appraisal or valuation services, actuarial services, management functions, human resources, investment advising, legal services, or expert services unrelated to the audit.99Id. § 210.2-01(c)(4)(i)–(x). Much of the intent behind this regulation is to reduce the types of conflicts discussed in Part II, in which auditors are incentivized to ignore errors in the audit in order to receive revenue for these non-audit services.100Paul Munter, The Importance of High Quality Independent Audits and Effective Audit Committee Oversight to High Quality Financial Reporting to Investors, U.S. Secs. & Exch. Comm’n (Oct. 26, 2021), https://www.sec.gov/news/statement/munter-audit-2021-10-26 [https://perma.cc/4P5E-UBX6]. Additionally, audit partners are required to rotate every five years if they are the lead partner on the engagement or every seven years otherwise, to avoid forming ties with clients that may impair independence.10117 C.F.R. § 210.2-01(c)(6)(i)(A)(1)–(2). Auditors are also not allowed to receive contingent fees or have partners compensated for any services other than audit services.102Id. § 210.2-01(c)(5). The audit engagement, typically including fees and non-audit services, must also be approved by the client’s audit committee to ensure independence.103Id. § 210.2-01(c)(7).

3.  Enforcement and Entanglement

The SEC’s enforcement mechanism for auditor independence violations is relatively straightforward. Under the SEC independence rules, the SEC can “censure a person or deny, temporarily or permanently, the privilege of appearing or practicing before [the SEC] in any way to any person who is found by the Commission after notice and opportunity for hearing”104Id. § 201.102(e)(1). to have “engaged in unethical or improper professional conduct.”105Id. § 201.102(e)(1)(ii). What is interesting about this mechanism is that it defers to “applicable professional standards” in determining whether there has been “unethical or improper professional conduct.”106Id. § 201.102(e)(1)(iv)(A), (e)(1)(ii). As a result, the SEC defers the determination of the standard for violations of independence standards to the body that sets the professional standards, which has often been interpreted, both by the SEC’s Administrative Law Judges and the federal courts, to be the AICPA.107For a discussion of cases that defer to AICPA auditor independence standards, see infra Part IV. This is one instance in which the auditor independence regulations are entangled between two different standard setters—the SEC and the AICPA.

One other instance in which the auditor independence rules are entangled between regulators is that although the SEC has delegated the authority to the PCAOB to set auditor independence standards, the SEC independence rules passed in the wake of Sarbanes-Oxley in 2003 are still effective.108O’Connor, supra note 49, at 565. Aside from some minor updates to debtor-creditor relationships passed in 2019, the SEC independence rules remained largely untouched until 2020,109Qualifications of Accountants, 85 Fed. Reg. 80508 (Dec. 11, 2020) (codified at 17 C.F.R. pt. 210). at which point the SEC, along with making minor updates to the independence rules, brought the PCAOB rules into harmony with the SEC rules where they conflicted, acknowledging that for several years there had been a period in which the SEC and PCAOB rules surrounding auditor independence were not consistent.110Id. This inconsistency is explored further in the sections below, including analyzing the frequency with which the courts used the SEC and PCAOB standards during this time, or the AICPA standards, which are defined in Part III.D.

C.  PCAOB Standards

The PCAOB sets forth “Ethics and Independence” standards for accounting firms and their associated persons.111Ethics & Independence, Pub. Co. Acct. Oversight Bd., https://pcaobus.org/oversight/

standards/ethics-independence-rules [https://perma.cc/WP6J-2LLD].
While several of these rules set forth additional requirements when compared to the SEC rules, several are similar or aligned with the SEC independence rules and AICPA standards.112See, e.g., Professional Standards, Rule 3521 (Pub. Co. Acct. Oversight Bd. 2006). First, accounting firms and their employees are required to “comply with all applicable auditing and related professional practice standards,” which implies that SEC and AICPA standards are binding.113Professional Standards, Rule 3100 (Pub. Co. Acct. Oversight Bd. 2003). However, in 2003, the PCAOB released a note to PCAOB Rule 3500T, which states that the “[PCAOB’s] Interim Independence Standards do not supersede the [SEC’s] auditor independence rules” and that in situations when the SEC Rules are more or less restrictive than the PCAOB rules, the more restrictive rule is to be followed.114Professional Standards, Rule 3500T (Pub. Co. Acct. Oversight Bd. 2003). Additionally, the PCAOB explicitly requires that accounting firms and their employees comply with AICPA Code of Professional Conduct Rules 101 and 102, including any interpretations and rulings under these rules.115Id. Rule 3500T(a), (b)(1). Further analysis of the AICPA rules will be provided in Section III.C.

The PCAOB independence rules extend beyond the SEC independence rules in three areas: (1) limiting auditors’ ability to provide audit clients with tax services,116Professional Standards, Rule 3522 (Pub. Co. Acct. Oversight Bd. 2006); Professional Standards, Rule 3523 (Pub. Co. Acct. Oversight Bd. 2006). (2) requiring auditors to communicate with the client Board’s audit committee about certain independence-related matters,117Professional Standards, Rule 3524 (Pub. Co. Acct. Oversight Bd. 2006); Professional Standards, Rule 3525 (Pub. Co. Acct. Oversight Bd. 2007); Professional Standards, Rule 3526 (Pub. Co. Acct. Oversight Bd. 2008). and (3) requiring auditors to submit a form to the PCAOB that summarizes audit hours by partner (“Form AP”).118Professional Standards, Rule 3211(a) (Pub. Co. Acct. Oversight Bd. 2016).

1.  Tax Services

Generally, the PCAOB rules extend beyond the SEC rules by maintaining that an accounting firm is not independent of its audit client if the firm provides “marketing, planning, or opining in favor of the tax treatment of” confidential transactions or aggressive tax position transactions, or if the firm provides “tax service to a person in a financial reporting oversight role” at the client, which generally prohibits providing tax advice or preparation services to management and finance employees of the audit client.119Professional Standards, Rule 3522 (Pub. Co. Acct. Oversight Bd. 2006); Id. Rule 3523.

2.  Audit Committee Communication

Further, the PCAOB goes beyond the SEC rules and requires accounting firms to seek audit committee pre-approval before the firms perform any permissible tax service or non-audit service related to internal controls.120Id. Rule 3524; Professional Standards, Rule 3525 (Pub. Co. Acct. Oversight Bd. 2007); Professional Standards, Rule 3526 (Pub. Co. Acct. Oversight Bd. 2008). Part of the intent of these rules is to present to the audit committee the ways that it might impair the accounting firm’s independence to provide both tax and non-audit services to the client.121See Professional Standards, Rule 3524(b) (Pub. Co. Acct. Oversight Bd. 2006); Professional Standards, Rule 3525(b) (Pub. Co. Acct. Oversight Bd. 2007). In line with this requirement, accounting firms are required to describe to the audit committees of their clients “all relationships between the registered public accounting firm . . . and the audit client or persons in financial reporting oversight roles at the potential audit client that . . . may reasonably be thought to bear on independence” both prior to beginning the audit and annually thereafter, including an annual affirmation of independence to the audit committee, which essentially requires accounting firms to disclose personal relationships that may impair independence.122Professional Standards, Rule 3526(a)(1) (Pub. Co. Acct. Oversight Bd. 2008); Professional Standards, Rule 3526(b)(2)–(3) (Pub. Co. Acct. Oversight Bd. 2008).

3.  Form AP Filing Requirements

Finally, the PCAOB requires that accounting firms file a “Form AP” with the PCAOB for each audit report it issues for a client.123Professional Standards, Rule 3211 (Pub. Co. Acct. Oversight Bd. 2016). The Form AP lists the lead engagement partner for the audit and notes the hours they have completed on the audit engagement, with the goal of providing users of financial statements information about the “independence of the specific individuals and firms that participate in the audit.”124Pub. Co. Acct. Oversight Bd., Supplemental Request for Comment: Rules to Require Disclosure of Certain Audit Participants on a New PCAOB Form A2–2 (2015), https://pcaob-assets.azureedge.net/pcaob-dev/docs/default-source/rulemaking/docket029/release_2015_

004.pdf [https://perma.cc/U3D4-2X49] .

D.  AICPA Standards

The AICPA promulgates what is the most rigorous standard of independence requirements for auditors, when compared to the SEC and PCOAB. The AICPA Code of Professional Conduct begins by setting forth a single independence rule: “A member in public practice shall be independent in the performance of professional services as required by standards promulgated by bodies designated by [the AICPA Governing] Council.”125Am. Inst. Certified Pub. Accts., Code of Professional Conduct, Rule 1.200.001.01. Stemming from this rule, the AICPA has then set forth hundreds of interpretations,126Id. Rule 1.200.001–1.298.010. which are binding on accounting firms and accountants performing “attest engagements,” which are any services to a client requiring independence from the client, including audits.127Am. Inst. Certified Pub. Accts., Plain English Guide to Independence 2 (2021), https://us.aicpa.org/content/dam/aicpa/interestareas/professionalethics/resources/tools/downloadabledocuments/plain-english-guide.pdf [https://perma.cc/N4AS-E3NG].

Given the complexity of the interpretations to the independence rule and the many scenarios effecting independence that they address, it would be impractical to describe all interpretations within this paper. However, to summarize, the interpretations to the independence rule cover a variety of situations regarding the independence of individual accountants, such as how to handle the employment of a family member at an audit client128Am. Inst. Certified Pub. Accts., Code of Professional Conduct, Rule 1.270.020.01–.03. and how to address whether an individual accountant’s financial investments in a client’s securities impair their independence.129Id. Rule 1.240.010.01–.03. The interpretations also cover more complex situations regarding the independence of accounting firms as a whole, such as how to maintain independence in situations in which a nonclient acquires a current client130Id. Rule 1.224.010.05–.08. or how “network firms” with multiple offices should each remain independent of each other’s clients.131Id. Rule 1.220.010.04.

In the absence of any relevant interpretation, accounting firms and accountants are expected to apply the AICPA’s “Conceptual Framework for Independence.”132Id. Rule 1.200.005.01. The Conceptual Framework for Independence requires that accountants evaluate whether a particular “relationship or circumstance” would lead a reasonable person “to conclude that there is a threat to . . . independence . . . that is not at an acceptable level.”133Id. Rule 1.210.010.01. The Conceptual Framework then lists potential “threats” to independence, such as holding an adverse interest from the client, advocating for the client, familiarity with the client, auditor participation in management of the client, self-interest, self-review, and undue influence by the client or a third party.134Id. Rule 1.210.010.10–.18. Following the potential threats, the Conceptual Framework identifies “safeguards” which can reduce threats to an acceptable level that will ensure independence, such as external review, competency requirements for professional licensing, analysis of an accounting firm’s revenue dependence on one client, and accounting firm policies for engagement quality control, such as external review.135Id. Rule 1.000.010.21–.23. By ensuring threats are low or nonexistent or by balancing them with safeguards, accounting firms and accountants can ensure compliance with independence standards under the Conceptual Framework in situations in which there is no authoritative interpretation set forth by the AICPA.136Id. Rule 1.210.010.07.

The AICPA also has a senior committee, the Auditing Standards Board, that sets forth Generally Accepted Auditing Standards (“GAAS”).137Clarified Statements on Auditing Standards, Am. Inst. Certified Pub. Accts. (Nov. 9, 2022), https://us.aicpa.org/research/standards/auditattest/clarifiedsas.html [https://perma.cc/7DN3-Q2CS]. GAAS sets forth specific testing requirements and procedures for auditors as they undertake audit engagements for clients.138Id. In addition to these audit testing requirements—which are largely technical in nature and outside of the scope of this paper—GAAS also sets for ethical requirements relating to audits of financial statements, stating that an “auditor must be independent of [a client] when performing an engagement in accordance with GAAS.”139Codification of Statements on Auditing Standards, AU-C § 200.15 (Am. Inst. of Certified Pub. Accts. 2012). GAAS then defers to the Code of Professional Conduct’s Conceptual Framework, which was discussed previously.140Id. § 200.17 (Am. Inst. of Certified Pub. Accts. 2012). GAAS is often referenced more frequently than the Code of Professional Conduct, as will be discussed in Part IV, because GAAS includes not only ethical requirements for auditors, but substantive guidance for how audits are to be conducted in practice.141Id.

Overall, there is significant overlap between the regulations set forth by the AICPA and the SEC and PCAOB.142See Plain English Guide to Independence, supra note 127. The AICPA standards set forth detailed guidance for accounting firms to ensure they comply with the broader standards set forth by the SEC and PCAOB. For example, while the SEC and PCAOB rules prevent auditors from receiving contingent fees from their clients, the AICPA rules state that same proposition and provide guidance that states contingent fees include finder’s fees, fees based on cost-savings achieved by the client, and exclude fees based on the results of judicial proceedings in tax matters.143Id. at 42. This level of detailed guidance means that practitioners often consult the AICPA standards, as they set forth a more restrictive set of guidelines and also a more informative set of interpretations that can be applied to specific circumstances. This delegation of substantive regulation to the private industry organization—AICPA—is also consistent with other self-regulatory models.

E.  Harmonization Efforts

In October 2020, the SEC issued updates to the auditor independence rules set forth in SEC Independence Rule 2-01.144Press Release, U.S. Secs. & Exch. Comm’n, SEC Updates Auditor Independence Rules (Oct. 16, 2020), https://www.sec.gov/news/press-release/2020-261 [https://perma.cc/N4AS-E3NG]. These updates covered a variety of miscellaneous matters in the SEC independence rules, including refining the definition of affiliates of audit clients, amending the definition of an “audit and professional engagement period,” excluding some student loans from causing independence violations, and addressing inadvertent violations of the independence rules due to mergers and acquisitions, among other matters.145Id.

As a result, stakeholders raised concerns that the PCAOB rules, particularly those related to affiliates of audit clients, such as subsidiaries, were no longer consistent with the SEC independence rules.146Pub. Co. Acct. Oversight Bd., PCAOB Release No. 2020-003, Amendments to PCAOB Interim Independence Standards and PCAOB Rules to Align with Amendments to Rule 2-01 of Regulation S-X 3 (Nov. 19, 2020), https://pcaob-assets.azureedge.net/pcaob-dev/docs/default-source/rulemaking/docket-047/2020-003-independence-final-rule.pdf [https://perma.cc/NW28-LB49]. In response, and to “provide greater regulatory certainty,” the PCAOB amended its rules to align with the SEC independence rule changes.147Id. Additionally, several of the PCAOB rules that needed realignment with the new SEC rules were those that the PCAOB adopted directly from the AICPA or had interpreted based on the AICPA rules.148Id. at 10. As a result, the AICPA issued a temporary policy statement as a stop-gap measure that stated to accounting firms and accountants that they would be considered in compliance with the AICPA Code of Professional Conduct if they complied with the updated SEC independence rules.149Temporary Policy Statement Related to Amendments of Rule 2-01 of Regulation S-X 4, Am. Inst. Certified Pub. Accts. (Dec. 21, 2020), https://us.aicpa.org/content/dam/aicpa/interestareas/
professionalethics/community/exposuredrafts/downloadabledocuments/2021/2021JanuaryOfficialReleaseTemporaryPolicyStatement.pdf [https://perma.cc/BJ8J-RRTT].
Soon thereafter, the AICPA put forth a proposal to change its rules and definitions to align with the SEC and PCAOB changes.150Am. Inst. Certified Pub. Accts., Proposed Revised Interpretations and Definition of Loans, Acquisitions, and Other Transactions 2–3 (2021), https://us.aicpa.org/content/dam/
aicpa/interestareas/professionalethics/community/exposuredrafts/downloadabledocuments/2021/2021-october-sec-loans-convergence.pdf [https://perma.cc/PR6C-9RJJ].
This proposal has not been adopted as of the time of this paper, but it seems likely that the AICPA will bring its standards into alignment with the current SEC and PCAOB independence rules.

Overall, this amendment waterfall that began with the SEC proposing changes in October 2020 that were then adopted by the PCAOB, and which still have not been adopted by the AICPA two years later, shows that the regulatory framework for auditor independence remains entangled. This entanglement may be occurring because the Sarbanes-Oxley Act created the PCAOB, which is allowed to pass audit regulations subject to the oversight of the SEC; and the Sarbanes-Oxley Act also allowed the SEC to become more involved in rule-setting for auditors, which had previously been handled entirely by the AICPA.15115 U.S.C. §§ 7211(a), 7233(a). While some argue that harmonization between these three entities is a worthy goal to disentangle regulations that are not consistent between the three entities,152See William D. Duhnke, Statement on Amendments to PCAOB Interim Independence Standards to Align with Amendments to Rule 2-01 of Regulation S-X, Pub. Co. Acct. Oversight Bd. (Nov. 19, 2020), https://pcaobus.org/news-events/speeches/speech-detail/statement-on-amendments-to-pcaob-interim-independence-standards-to-align-with-amendments-to-rule-2-01-of-regulation-s-x [https:
//perma.cc/7FDK-CWTT].
one might also consider whether harmonization is the answer, or rather whether simplification is a better goal, to avoid having to involve three regulatory agencies in rule changes, in order to ensure clear standards for accounting firms, clients, and stakeholders in the market. The following section of the paper will explore whether one regulatory entity is dominant over the others, and whether this agency should be favored for simplification that centers around focusing auditor independence on this agency and its rules exclusively, rather than the standards set across all three entities.

IV.  CASE STUDY

As discussed, the following case study will assess published federal court opinions in which auditor independence was at issue in a civil litigation to determine whether SEC, PCAOB, or AICPA standards were used in the courts’ reasoning.153For a full listing of cases reviewed, see infra APPENDIX. While a study of administrative law decisions from the SEC was considered, federal court decisions were deemed to be a more relevant indicator of which body of regulation is used in deciding civil matters because SEC administrative law decisions rely almost exclusively on a single, broad rule, SEC Rule 102(e)(1)(ii), which “censure[s] a person . . . after finding that a person engaged in improper professional conduct.”154See, e.g., Order Instituting Public Administrative and Cease and Desist Proceedings, In the Matter of Alan C. Greenwell, CPA, U.S. Secs. & Exch. Comm’n (Dec. 10, 2021), https://www.sec.gov/litigation/admin/2021/34-93750.pdf [https://perma.cc/44M7-22PR].

There were fifteen cases decided by federal courts in which auditor independence was at issue in civil litigation, and which were used to comprise the population for this case study.155For a full listing of cases reviewed, see infra APPENDIX. The population of cases includes only cases in which auditors were performing financial statement audits, and therefore excludes government and internal audit services, as these non-financial statement audits are typically not part of the debate over auditor independence because they involve different monetary incentives, risks for auditors and investors, and different auditing standards. Additionally, cases were only observed after May 6, 2003, which was the effective date of the Sarbanes-Oxley legislation, given that Sarbanes Oxley completely changed the landscape of auditor independence, giving the SEC broad authority to issue rulemaking in this area and effectively creating the PCAOB.156Strengthening the Commission’s Requirements Regarding Auditor Independence, Release No. 33-8183, 68 Fed. Reg. 6005 (codified as 17 C.F.R. §§ 210, 240, 249, 274 (2003)). Finally, cases that relied on state regulations over auditor licensing and independence were excluded, as they do not address the relevant issue of federal regulatory structure.157There was only a single case that relied on state professional licensing requirements, Rahl v. Bande, 328 B.R. 387 (S.D.N.Y. 2005). Given that this analysis is focused on federal regulation and this case is an outlier in that it is the only federal case that relies on state professional licensing regulation in its reasoning, it has been excluded from the population.

This case study aims to examine which body of regulatory law federal courts rely on in making determinations over whether auditors have breached their independence obligations. Each of the AICPA, SEC, and PCAOB have their own enforcement mechanisms to sanction auditors who do not adhere to independence requirements, and each enforcement division uses their own regulation as well as occasionally relies on the regulation of the other entities to sanction auditors.158See Professional Standards, Rules 5000­5501 (Pub. Co. Acct. Oversight Bd. 2004); Ethics Enforcement, Am. Inst. Certified Pub. Accts., https://us.aicpa.org/interestareas/
professionalethics/resources/ethicsenforcement [https://perma.cc/YDV6-X4S8]; Accounting and Auditing Enforcement Releases, U.S. Secs. & Exch. Comm’n, https://www.sec.gov/
divisions/enforce/friactions.htm [https://perma.cc/YDV6-X4S8].
However, the civil courts do not have a requirement to adhere to the regulations of any particular standard-setter under the requirements of Sarbanes-Oxley or the Securities and Exchange Act of 1934. Given this lack of constraints, the standard used by the federal courts in determining independence violations has not been examined and is ripe for analysis to determine whether one set of standards (that is, those of the AICPA, SEC, or PCAOB) is preferred over the others. Therefore, this study will examine all fifteen federal court decisions to determine what standards have been used by the courts in the area of auditor independence as well as whether the result was in favor of the auditor or against the auditor in each scenario.

There are three major areas in which auditor independence has been examined by the federal courts. First, individuals who have been sanctioned by the SEC for violations of auditor independence standards can appeal to the federal courts for a review of the SEC’s decisions under a broad “abuse of discretion” standard to argue that the agency acted arbitrarily and capriciously in a way that requires the SEC’s decision to be overturned under 5 U.S.C. 706(2)(A).159Ponce v. U.S. Secs. & Exch. Comm’n, 345 F.3d 722, 728–29 (9th Cir. 2003).

Next, auditor independence is frequently at issue in Securities and Exchange Act Section 10(b) and SEC Rule 10b-5 claims, which are often brought as class-actions.160See, e.g., In re WorldCom, Inc. Sec. Litig., 352 F. Supp. 2d 472, 497 (S.D.N.Y. 2005). To prevail on these claims, plaintiffs must prove scienter, meaning that the defendant employed a “device, scheme, or artifice to defraud,” in addition to proving the elements of a material misstatement or omission, on which the plaintiff relied, and was the proximate cause of the plaintiff’s loss.16117 C.F.R. § 240.10b-5(a)–(c); 4 James D. Cox & Thomas Lee Hazen, Treatise on the Law of Corporations § 27:19 (3d ed. 2022). Plaintiffs often allege that auditors’ violations of independence rules are evidence of scienter for purposes of satisfying this element to bring a successful 10(b) or 10b-5 claim.162See, e.g., In re WorldCom, Inc., 352 F. Supp. 2d at 497. However, this practice has been complicated by the heightened pleading requirements for Section 10(b) and Rule 10b-5 claims established in the Private Securities Litigation Reform Act of 1995 (“PSLRA”), which was passed in part to reduce the number of non-meritorious securities class action claims raised by plaintiffs and requires that plaintiffs “state with particularity facts giving rise to a strong inference that the defendant acted with the required state of mind.”16315 U.S.C. § 78u-4(b). The PSLRA has resulted in different pleading standards for scienter among and within circuits, however, many courts find “scienter [is] plead with particularity by facts supporting a ‘motive or opportunity’ to commit fraud.” 164Cox & Hazen, supra note 161. As will be discussed in Section IV.B, this heightened pleading standard has reduced the circumstances in which courts have viewed violations of auditor independence rules to be sufficient to show scienter.

Finally, plaintiffs also bring state law claims—including negligent misrepresentation, professional negligence, and fraud suits—against auditors who have violated auditor independence standards, using the alleged violations of these standards as de facto evidence of a breach of duty.165See, e.g., New Jersey v. Sprint Corp., 314 F. Supp. 2d 1119, 1126, 1134 (D. Kan. 2004); In re Parmalat Sec. Litig., 501 F. Supp. 2d 560, 566 (S.D.N.Y. 2007); Newby v. Enron Corp. (In re Enron Corp. Secs., Derivative & ERISA Litig.), 762 F. Supp. 2d 942, 954 (S.D. Tex. 2010). In one instance, a plaintiff also attempted to bring a state law breach of fiduciary duty claim against an auditor who allegedly did not comply with auditor independence standards.166In re SmarTalk Teleservices, Inc. Secs. Litig., 487 F. Supp. 2d 928, 931 (S.D. Ohio 2007).

Given the heterogeneity of the claims within these cases, this paper examines the cases within these three groups—reviews of SEC administrative decisions, federal securities law claims, and state law tort claims—to identify developments in the case law regarding auditor independence and to examine when courts apply SEC, PCAOB, or AICPA standards in determining whether there has been an independence violation sufficient to warrant a judgment against auditors.

A.  Appeals of SEC Administrative Decisions

In Ponce v. SEC, the Ninth Circuit reviewed an appeal from a decision that the SEC made to bar a plaintiff accountant from practice, and the court acknowledged that as part of the SEC’s decision, the accountant had been held in violation of SEC Independence Rule 102(e)(1)(ii), which means he engaged in “improper professional conduct.”167Ponce v. U.S. Secs. & Exch. Comm’n, 345 F.3d 722, 739 (9th Cir. 2003). In order to determine whether there was truly “improper professional conduct,” the court turned to AICPA standards and ruled that the auditor failed to maintain his independence because he allowed his clients to run up a substantial balance of unpaid fees, which under AICPA guidance, resulted in a presumed lack of independence because the AICPA standards set forth that “independence is considered to be impaired if fees for all professional services rendered for prior years are not collected before the issuance of the member’s report for the current year.”168Id. at 728.

Similarly, in Dearlove v. SEC, the Court of Appeals for the District of Columbia Circuit used the same approach as Ponce, albeit six years after the decision169Dearlove v. U.S. Secs. & Exch. Comm’n, 573 F.3d 801, 804 (D.C. Cir. 2009). and six years after the implementation of Sarbanes-Oxley, including its resulting reform of SEC independence rules and the formation of the PCAOB. In Dearlove, the court concluded that “the appropriate standard of care . . . is supplied by . . . GAAS” when reviewing whether the SEC had abused its discretion in determining that an accountant violated SEC Independence Rule 102(e)(1)(ii) by failing to maintain independence from their audit client.170Id.

Ponce and Dearlove are indicators of how the federal courts have given credibility to the AICPA standards, including GAAS, in determining whether there has been a violation of auditor independence. The court’s opinion in Dearlove references SEC Independence Rule 102(e)(1)(ii), but it does so only to note that the SEC rule states “improper professional conduct” will warrant sanctions, before deferring to the AICPA GAAS standards to assess what improper conduct is.171Id. at 803–804. The court also stated that “the SEC need not establish a standard of care separate from the GAAS in order to give meaning to” what SEC Independence Rule 102(e)(1)(iv)(B)(2) describes as “unreasonable conduct,” showing further deference to AICPA standards.172Id. at 805–06. The fact that this decision came more than six years after the implementation of Sarbanes-Oxley, when the court had the ability to reference the updated SEC independence rules or PCAOB rules, but chose not to, shows even more significant reliance on the standards set by the AICPA.

B.  Determinations of Scienter in Securities Claims

Similarly, in In re WorldCom, Inc. Securities Litigation, defendants moved for summary judgment on the matter of whether Arthur Andersen, the auditors at the helm of both the Enron, and in this case, WorldCom accounting scandals, had the requisite scienter to be in violation of Section 10(b) of the Securities and Exchange Act and SEC Rule 10b-5 because they allegedly recklessly issued false audit opinions.173In re WorldCom, Inc., 352 F. Supp. 2d at 494–95. The plaintiff alleged that violations of the AICPA’s GAAS were sufficient to prove scienter, even under the heightened pleading standards required by the PSLRA, which required the plaintiffs to plead recklessness in order to avoid their claim being dismissed.174Id. at 495, 497. The court recognized the importance of violations of GAAS in proving scienter, but ultimately denied summary judgment due to conflicting expert reports on whether GAAS was violated.175Id. at 499–500. Although the claim survived the motion for summary judgment due to unresolved questions of fact, this case was another high profile example of the federal courts giving credence to AICPA standards in determining whether there was auditor wrongdoing.176Id.

In re WorldCom, Inc. Securities Litigation also included a Securities Act claim, in which the plaintiffs alleged that Arthur Andersen was in violation of Section 11 of the Securities Act, which states that a “preparing or certifying accountant . . . may be liable ‘if any part of the registration statement . . . contained an untrue statement of a material fact.’ ”177Id. at 490–91 (quoting 15 U.S.C. § 77k(a)). Arthur Andersen attempted to assert a due diligence defense, in which it claimed that it “had, after reasonable investigation, reasonable ground to believe and did believe, at the time . . . the registration statement became effective, that the statements therein were true,” and then moved for summary judgment.178Id. at 491–92. In deciding whether summary judgment was appropriate on this issue, the court issued an even stronger affirmance of the relevance of AICPA standards, concluding that a “reasonable investigation” that would support a due diligence defense, like that raised by Arthur Andersen, must be a “GAAS-compliant audit.”179Id. at 492. Because the plaintiff presented sufficient evidence to rebut the argument that the audit was “GAAS-compliant,” Arthur Andersen’s motion for summary judgment was denied. While this second issue is not directly related to auditor independence rules, it shows the courts’ general deference to the AICPA’s GAAS.

However, not all courts have agreed with the Southern District of New York’s decision in WorldCom, which may be in part due to differing interpretations of the heightened pleading requirements of the PSLRA. In In re Cardinal Health, Inc. Securities Litigations, the Southern District Court of Ohio ruled that Ernst & Young’s failure to adhere to the AICPA’s GAAS requirements for auditor independence did not establish the requisite scienter for a plaintiff’s claim to survive Ernst & Young’s motion to dismiss on a Rule 10b-5 claim.180In re Cardinal Health, Inc. Sec. Litigs., 426 F. Supp. 2d 688, 697–98 (S.D. Ohio 2006). The Court reasoned that while recklessness is generally sufficient to meet the pleading standard under the PSLRA, claims brought against auditors were subject to the even more heightened pleading standard of “a mental state ‘so culpable that it approximate[s] an actual intent to aid in the fraud being perpetrated by the audited company,’ ” which was not met in this case based merely on the alleged failure of Ernst & Young to adhere to AICPA GAAS requirements.181Id. at 763 (quoting Fidel v. Farley, 392 F.3d 220, 227 (6th Cir. 2004)) (internal quotations omitted).

Further, the court opined that an auditor’s past sanctions in SEC administrative proceedings were insufficient to prove scienter.182Id. at 778–79. In this case, the court also noted that SEC administrative decisions were not dispositive in determining scienter for Rule 10b-5 claims, perhaps suggesting that judicial interpretation of independence violations supersedes determinations by regulatory bodies.183See id. This is in line with existing administrative law doctrines that do not require federal courts to defer to the SEC’s interpretations of the Exchange Act.184U.S. Secs. & Exch. Comm’n v. McCarthy, 322 F.3d 650, 654 (9th Cir. 2003).

Similarly, in In re Royal Ahold N.V. Securities and ERISA Litigation, the United States District Court for the District of Maryland ruled that alleged violations of AICPA’s GAAS standards on independence—in which the only allegations from the plaintiff were that auditor independence was impaired due to the auditor providing audit and non-audit services—were not sufficient to establish scienter for a Rule 10b-5 claim against an auditor.185In re Royal Ahold N.V. Sec. & ERISA Litig., 351 F. Supp. 2d 334, 390–92 (D. Md. 2004). However, the court did note that violations of AICPA’s GAAS can be sufficient to plead scienter when they are coupled with allegations that show that “the nature of the violations of those violations was such that scienter is properly inferred.”186Id. at 386. Likewise, in New Jersey v. Sprint Corp., a group of class action plaintiffs brought Rule 10b-5 claims against Sprint Corporation and Ernst & Young for filing false and misleading registration statements, prospectus supplements, and proxy statements that did not disclose that the company had considered dismissing their auditor, Ernst & Young, and that there were conflicts between executives at the company and the auditor, which resulted in auditor independence violations under AICPA’s GAAS.187New Jersey v. Sprint Corp., 314 F. Supp. 2d 1119, 1126, 1123, 1126, 1134 (D. Kan. 2004). The court relied on the AICPA’s GAAS to assess independence, noting that the plaintiffs did not plead sufficient facts to show that Ernst & Young lacked the independence in “mental attitude” required by GAAS, and even if sufficient facts were pled to show that Ernst & Young violated GAAS, this would not be sufficient under the PSLRA to establish scienter because the standard for scienter is recklessness, that is “so obvious that the defendant must have been aware of it” which goes beyond a mere violation of GAAS. 188Id. at 1134–35, 1147–48. Therefore, the motion to dismiss was granted in favor of Ernst & Young.189Id. at 1149.

While the varied interpretations in different jurisdictions over whether violations of the AICPA’s GAAS standards is sufficient to plead scienter persists, some courts have ruled that merely “articulating violations of GAAS and GAAP alone is insufficient” to satisfy the element of scienter under the PSLRA’s heightened pleading requirements and instead imposed additional requirements to plead scienter through case law.190Grand Lodge of Pa. v. Peters, 550 F. Supp. 2d 1363, 1372 (M.D. Fla. 2008). For example, the court in Grand Lodge of Pennsylvania v. Peters determined that in order to prove scienter, violations of GAAS must be accompanied by “red flags” that would put a reasonable auditor on notice that their client was committing fraud.191Id. at 1372. Therefore, the plaintiff’s allegations in this case that the auditor was conflicted by providing consulting services to the client in violation of GAAS were insufficient to establish scienter for a Rule 10b-5 claim.192Id at 1372–73. Additionally, the court in In re Williams Securities Litigation ruled that “GAAS violations must be coupled with evidence that the violations were the result of the auditor’s fraudulent intent to mislead investors,” in order to have a sufficient pleading of scienter.193In re Williams Sec. Litig., 496 F. Supp. 2d 1195, 1289 (N.D. Okla. 2007). This supports the idea that some courts grant credibility to the AICPA standards, however, they impose additional burdens on plaintiffs that are developed through case law.

This case law has developed in the years since the passage of the PSLRA. The Ninth Circuit Court of Appeals summarized the development of the additional requirements to prove scienter, other than GAAS violations, in New Mexico State Investment Council v. Ernst & Young LLP.194N.M. State Inv. Council v. Ernst & Young LLP, 641 F.3d 1089, 1097–98 (9th Cir. 2011). The court ruled that failing to “maintain independence in mental attitude during an audit,” in violation of GAAS and PCAOB standards, is not sufficient to prove scienter.195Id. at 1097. Rather, there should be “red flags” that a reasonable auditor would have investigated as well as a showing that there were violations that amount to more than “alleging a poor audit.”196Id. at 1098. New Mexico State Investment Council indicates how the courts’ reliance solely on AICPA standards has lessened slightly over the years, as the burden to meet additional case law requirements for scienter has increased due to the passage of the PSLRA and its heighted pleading requirements, which require that plaintiffs “state with particularity facts giving rise to a strong inference that the defendant acted with the required state of mind.”19715 U.S.C. § 78u-4(b)(2)(A).

While the discussion above indicates that courts have largely relied on violations of AICPA’s GAAS and case law requirements in determining whether auditor independence violations are sufficient to show scienter, other courts have discussed Sarbanes-Oxley in determining whether an auditor independence violation is sufficient to show scienter as an element of a Rule 10b-5 violation.198Brody v. Stone & Webster, Inc. (In re Stone & Webster, Inc., Sec. Litig.), 414 F.3d 187, 215 (1st Cir. 2005). In Brody v. Stone Webster, Inc., the First Circuit Court of Appeals examined whether scienter could be presumed on the part of PwC in connection with a 10b-5 claim because PwC allegedly turned a blind eye toward accounting irregularities to protect its accounting and consulting revenue from a client.199Id. The court determined that “turn[ing] a bind [sic] eye” to misleading accounting for a “profit motive” may have been a rationale for the passage of Sarbanes Oxley, but it is not enough in itself to prove scienter sufficient for a valid Rule 10b-5 claim under the PSLRA without specific allegations that the auditor ignored “red-flags” that were signs of fraud.200Id.

Courts have also been reluctant to find that scienter has been sufficiently pled in accordance with the PSLRA when plaintiffs allege auditor independence violations on the basis of general policy arguments against auditors depending on fees from clients.201See In re ArthroCare Corp. Secs. Litig., 726 F. Supp. 2d 696, 733 (W.D. Tex. 2010). In In re ArthoCare Corporation Securities Litigation, plaintiffs alleged that PwC was not independent in its audit because the firm had a longstanding relationship with its client and was dependent on the client’s audit fees.202Id. The United States District Court for the Western District of Texas found that general allegations based on a perceived lack of independence or violations of GAAS due to fee dependence or long-standing relationships, which are allegedly against public policy, are not sufficient to establish scienter on the part of auditors under the PSLRA.203Id. Similarly, in Ley v. Visteon Corporation, the Sixth Circuit determined that an allegation that Ernst & Young was not independent during an audit because it sought to preserve revenue from a client by not pointing out the client’s alleged accounting irregularities was not sufficient to plead scienter on the part of the auditors under the PSLRA’s requirements because it was merely an allegation of a “motive.”204Ley v. Visteon Corp., 543 F.3d 801, 815 (6th Cir. 2008).

On the whole, these cases show a broad trend of the courts giving credibility to the AICPA’s standards, as compared to SEC or PCAOB standards. As the case law has developed, violation of AICPA standards has been shown as one of the avenues plaintiffs can use to establish scienter for Securities and Exchange Act Section 10(b) and SEC Rule 10b-5 claims, when additional case law requirements are met. Notably, neither violations of the PCAOB independence rules nor the SEC independence rules have been used by the federal courts to assess whether there has been scienter for the purposes of a Section 10(b) or Rule 10b-5 claim that has been sufficiently pled in accordance with the PSLRA. Instead, the courts have relied on the AICPA rules as one factor for establishing scienter, and then have developed additional case law standards, such as finding “red flags” and alleging more than a poor audit, in order for plaintiffs to prevail on Section 10(b) and Rule 10b-5 claims against auditors. This focus on the AICPA rules is aligned with the deference given by the courts to AICPA standards in appeals of SEC enforcement decisions, discussed in Section IV.A.

C.  Fraud, Negligence, and Other State Law Causes of Action

In state law actions for fraud, courts have required a heightened standard for liability that extends beyond a violation of the AICPA’s GAAS independence standards in order to hold auditors liable. For example, in In re Parmalat Securities Litigation, plaintiffs pled that Deloitte aided an audit client with common law fraud.205In re Parmalat Sec. Litig., 501 F. Supp. 2d 560, 566 (S.D.N.Y. 2007). Given that this was a state law cause of action, the court applied the New York common law requirements for fraud, requiring the plaintiff to prove “that the defendant (1) made a material, false statement; (2) knowing that the representation was false; (3) acting with intent to defraud; and that plaintiff (4) reasonably relied on the false representation; and (5) suffered damage proximately caused by the defendant’s actions.”206Morris v. Castle Rock Ent., Inc., 246 F. Supp. 2d 290, 296 (S.D.N.Y. 2003). The court focused primarily on the intent to defraud and ruled that an auditor’s violation of GAAS does not by itself show intent, however “an auditor’s decision to take on non-audit work that threatens to compromise its duty of independence gives rise to a strong inference of . . . [fraudulent] intent . . . when . . . the auditor has a ‘direct stake’ in the alleged fraud.”207In re Parmalat, 501 F. Supp. 2d at 583–84. This heightened standard to show the intent element for common law fraud in New York, which includes not only a GAAS violation but also an auditor’s direct stake in the fraud, is in some ways analogous to the heightened standard to prove scienter on the part of auditors, as discussed in Section IV.B, because in both instances the courts have recognized a violation of the AICPA’s GAAS standards as helpful in showing intent, in addition to adding case law requirements to plead a valid claim.

However, some courts have not found it necessary to analyze auditor independence regulation when determining whether auditors were negligent in their review of their clients’ financial statements. In a consolidated action, insurance companies brought state law claims against Enron, Enron management, and Enron’s auditors, Arthur Andersen, in the wake of the Enron accounting scandal and resulting collapse.208Newby v. Enron Corp. (In re Enron Corp. Secs., Derivative & ERISA Litig.), 762 F. Supp. 2d 942, 954–55 (S.D. Tex. 2010). The plaintiffs specifically alleged that Arthur Andersen made negligent misrepresentations to investors and committed common law fraud in violation of Texas law.209Id. at 1021–22. In connection with the negligent misrepresentation claim, the plaintiffs were required to show (1) the defendant provided information (2) that was false, (3) the defendant did not exercise reasonable care or competence in obtaining or communicating the information, (4) the plaintiff justifiably relied on the information, and (5) the plaintiff suffered loss by justifiably relying on the information.210Id. at 980. In connection with the common law fraud claim under Texas law, the plaintiffs were required to show “(1) a material representation was made; (2) the representation was false; (3) when the representation was made, the speaker knew it was false or made it recklessly . . . (4) the representation was made with the intention that it be acted upon by the other party; (5) the party actually and justifiably acted in reliance upon the representation; and (6) the party suffered injury.”211Id. at 966. While the plaintiffs alleged that Arthur Andersen’s violations of AICPA’s GAAS standards showed the firm lacked independence, which they claimed was sufficient to show that Arthur Andersen did not meet the standard of care and competence as required by the negligent misrepresentation claim,212Id. at 1004. and that the violation of GAAS showed that Arthur Andersen made false representations about its independence sufficient for a common law fraud claim,213Id. at 1003–04. the court did not reach these issues, instead dismissing both claims because the plaintiffs could not show they relied on Arthur Andersen’s representations.214Id. at 1021–22.

Another area in which courts have assessed whether auditor independence gives rise to liability is in the area of state law claims for breach of fiduciary duties. In In re SmarTalk Teleservices, Inc. Securities Litigation, a trustee argued that PwC exceeded its normal role as an independent auditor, as defined by the AICPA’s GAAS, and therefore PwC owed a trustee fiduciary duties that the firm then breached by providing inadequate accounting and audit services.215In re SmarTalk Teleservices, Inc. Secs. Litig., 487 F. Supp. 2d 928, 931 (S.D. Ohio 2007). While the court did not reach the issue of whether there was a valid cause of action for breach of fiduciary duty, the court determined that whether PwC violated GAAS standards for independence was a genuine issue of material fact and denied the auditor’s motion for summary judgment on the breach of fiduciary duty claim, providing yet another example of the federal courts giving credence to the AICPA standards in determining auditor liability.216Id. at 935.

V.  IMPLICATIONS OF STUDY ON HARMONIZATION OR SIMPLIFICATION OBJECTIVES

In all three areas of the law covered by the case study, including appeals of SEC enforcement decisions, federal securities claims, and state law claims, the federal courts have preferred to use the AICPA standards, including GAAS, in their decision-making over auditor independence. In many instances when AICPA standards were used in federal securities actions, case law requirements were also supplied to determine whether there was sufficient evidence presented to plead scienter. But, notably, there were no cases found that apply PCAOB or SEC auditor independence rules.

On the whole, this paper does not aim to provide a normative proposal for auditor independence regulation. However, the case study presented in Part IV can shed light on the process of harmonization, which, as discussed in Section III.E, involves the SEC, PCAOB, and AICPA each having to change their rules regarding auditor independence each time one of the other entities changes their independence rules, in order to ensure that the rules are not in conflict. Given that the case study indicates that AICPA standards are preferred by the federal courts, along with case law, in determining whether auditor independence has been violated, there is an argument to be made that the power to regulate auditor independence should be simplified into one regulatory framework run by a single entity—the AICPA. Under this proposal of “simplification,” the regulations set by each entity would not need to be harmonized each time one entity makes a change in auditor independence rules. Rather, given that the federal courts rely on AICPA standards, the substantive rule-making authority would be given to the AICPA alone. This is because the current framework has two entities, the SEC and PCAOB, whose frameworks are rarely applied in the federal courts or outside of internal enforcement actions and investigations.

As discussed in Part III, the AICPA, PCAOB, and SEC are involved in a self-regulatory model, in which the SEC provides oversight over the PCAOB, and the SEC and PCAOB have concurrent jurisdiction to set auditor independence standards as a federal regulatory agency and as a non-profit corporation subject to the oversight of the SEC, respectively. In this self-regulatory model, the SEC also provides oversight over the AICPA, which is a private industry organization run by accountants and which sets substantive auditor independence regulation. This self-regulatory model could be simplified to reflect other self-regulatory structures in United States financial services regulation so that rule-making authority is deferred entirely to the AICPA, with oversight by the SEC, in order to simplify the regulatory framework and reduce conflicts between the independence rules set by the AICPA, SEC, and PCAOB. This would be similar to the model between the SEC and FINRA, in which the SEC allows FINRA to set rules for national securities exchanges and provides oversight over that rulemaking, rather than having the SEC set its own detailed regulation over exchanges. Given that the AICPA sets forth the most comprehensive rulemaking and interpretations of auditor independence standards, and the federal courts rely on these standards, simplifying the regulatory framework for auditor independence by deferring to the AICPA seems like a possible solution.

Overall, the relative costs and benefits of self-regulation and the outsourcing of rulemaking to private industry are beyond the scope of this paper; however, the following arguments are meant to provide a brief summary of why simplification of rule-making authority regarding auditor independence regulations by giving authority to the AICPA and oversight to the SEC may or may not be beneficial.

There are valid reasons to argue against simplification that comes in the form of allowing the AICPA to be the only standard-setter in the area of auditor independence. Several critics have pointed out that self-regulation was one of the issues at the forefront of the Enron collapse and resulting scandal, and that the accounting profession needs an external regulator.217U.S. Gov’t Accountability Off., GAO-02-411, The Accounting Profession: Status of Panel on Audit Effectiveness Recommendations to Enhance the Self-Regulatory System 1 (2002); see also Reed Abelson & Jonathan D. Glater, Enron’s Many Strands: The Auditors; Who’s Keeping the Accountants Accountable, N.Y. Times (Jan. 15, 2002), https://www.nytimes.com/
2002/01/15/business/enron-s-collapse-the-auditors-who-s-keeping-the-accountants-accountable.html [https://perma.cc/D6DH-V59V].
However, others have argued that it was not self-regulation, but market failures and misaligned incentives over reputational costs that caused the accounting scandals during the early 2000s.218Coffee, supra note 14, at 1420–21. Further, self-regulatory organizations, such as FINRA, have successfully provided guidance to their respective stakeholders, with some oversight from the SEC, indicating that self-regulatory organizations with some administrative oversight can be successful.219Luis A. Aguilar, The Need for Robust SEC Oversight of SROs, Harv. L. Sch. F. Corp. Governance (May 9, 2013), https://corpgov.law.harvard.edu/2013/05/09/the-need-for-robust-sec-oversight-of-sros [https://perma.cc/D6DH-V59V].

Additionally, critics may argue that the AICPA relies on the SEC and PCAOB enforcement practices in addition to running its own enforcement program,220Ethics Enforcement, supra note 158. and to split up the enforcement and regulation practices could pose problems. However, given that the current system splits the enforcement burden between the SEC, PCAOB, and AICPA, and each uses violations of the other’s regulations to bring sanctions, consolidating the regulations into one body would not have to change this framework.

CONCLUSION

This paper reserves judgment on the relative merits of self-regulation and instead notes that the current regulatory harmonization effort is not the only solution to disentangle the regulatory framework for auditor independence. Instead, this paper poses a new potential solution—simplification—to the problem of unwinding the tangled regulatory framework of auditor independence to promote efficiency in rulemaking and clarity for stakeholder accounting firms, regulators, and clients.

Given that the courts frequently defer to AICPA auditor independence standards—along with case law requirements for pleading federal securities law violations—rather than SEC and PCAOB standards, and having three regulatory frameworks that need to be continuously updated to align with each other is complex and costly, simplification is a worthy goal. However, it is just one solution of many. As the SEC, PCAOB, and AICPA continue to pursue harmonization,221Press Release, U.S. Secs. & Exch. Comm’n, SEC Updates Auditor Independence Rules (Oct. 16, 2020), https://www.sec.gov/news/press-release/2020-261 [https://perma.cc/4Q63-BQEW]. it is worth considering whether other alternative approaches to auditor independence regulation, such as simplification, exist.

APPENDIX

Appendix

Case Name and Citation

Procedural Posture

Body of Law Applied

Result in Favor of Auditor?

1

Ponce v. SEC, 345 F.3d 722 (9th Cir. 2003).

Appeal of Administrative Decision

AICPA and SEC

No

2

Dearlove v. SEC, 573 F.3d 801 (D.C. Cir. 2009).

Appeal of Administrative Decision

AICPA and SEC

No

3

New Jersey v. Sprint Corp., 314 F. Supp. 2d 1119 (D. Kan. 2004).

Motion to Dismiss

AICPA

Yes

4

Newby v. Enron Corp. (In re Enron Corp. Secs., Derivative & ERISA Litig.), 762 F. Supp. 2d 942 (S.D. Tex. 2010).

Motion to Dismiss

AICPAa

Yes

5

In re WorldCom, Inc. Sec. Litig., 352 F. Supp. 2d 472 (S.D.N.Y. 2005).

Motion for Summary Judgment

AICPA and SEC

No, on Securities Act Claim.

Yes, on SEC Rule 10b-5 claim.

6

In re Cardinal Health, Inc. Sec. Litigs., 426 F. Supp. 2d 688 (S.D. Ohio 2006).

Motion to Dismiss

AICPA

Yes

7

In re Royal Ahold N.V. Sec. & ERISA Litig., 351 F. Supp. 2d 334 (D. Md. 2004).

Motion to Dismiss

AICPA

Yes

8

Brody v. Stone & Webster, Inc. (In re Stone & Webster, Inc., Sec. Litig.), 414 F.3d 187 (1st Cir. 2005).

Motion to Dismiss

Sarbanes-Oxley

Yes

9

Ley v. Visteon Corp., 543 F.3d 801 (6th Cir. 2008).

Motion to Dismiss

AICPA

Yes

10

Grand Lodge of PA v. Peters, 550 F. Supp. 2d 1363 (M.D. Fla. 2008).

 

Motion to Dismiss

AICPA

Yes

11

In re Williams Sec. Litig., 496 F. Supp. 2d 1195 (N.D. Okla. 2007).

Motion for Summary Judgment

AICPA

Yes

12

N.M. State Inv. Council v. Ernst & Young LLP, 641 F.3d 1089 (9th Cir. 2011).

Motion for Summary Judgment

AICPA

Yes

13

In re Parmalat Sec. Litig., 501 F. Supp. 2d 560 (S.D.N.Y. 2007).

Motion to Dismiss

AICPA

Yes

14

In re ArthroCare Corp. Secs. Litig., 726 F. Supp. 2d 696 (W.D. Tex. 2010).

Motion to Dismiss

AICPA

Yes

15

In re SmarTalk Teleservices, Inc. Secs. Litig., 487 F. Supp. 2d 928 (S.D. Ohio 2007).

Motion for Summary Judgment

AICPA

No

Note:  The plaintiffs pled a violation of AICPA standards, but the court did not reach the issue in this case before making its final determination.

97 S. Cal. L. Rev. 495

Download

* Executive Articles Editor, Southern California Law Review, Volume 97; J.D. Candidate, University of Southern California Gould School of Law, 2024; B.S., B.A., Boston College, 2017. Many thanks to Professor Jonathan Barnett for his feedback and guidance, as well as to the editors of the Southern California Law Review for their thoughtful suggestions. All mistakes are my own.

Corporate Social Responsibility Through Shareholder Governance

New approaches to corporate purpose have emerged in recent years that hold out the promise of addressing concerns about corporate social responsibility (“CSR”) through shareholder governance, rather than in spite of it. The seminal such approach—enlightened shareholder value—posits that treating other stakeholders well can ultimately redound to long-term shareholder value. However, two more recent proposals reconceptualize shareholder interests in more holistic ways and urge that it is shareholders’ welfare, not shareholder value per se, that managers should pursue. In particular, the “shareholder social preferences” view incorporates into the corporate objective the degree to which the firm’s operations align with the social views of shareholders. The “portfolio value maximization view,” in contrast, argues that corporate fiduciaries should maximize the value of diversified shareholders’ portfolios by considering the externalities of the firm’s operations on those portfolios.

Shifting to shareholder welfare as the corporate objective, however, would do little to improve corporate conduct and would entail substantial costs. The social preferences of shareholders are conflicted, muted, and often prefer less protection of stakeholder interests than provided by law. Shareholders’ portfolio value captures only a small portion of the externalities like pollution that its proponents hope to address and risks motivating anticompetitive conduct. And neither corporate managers nor shareholders would have the information and incentives needed to pursue these additional shareholder welfare considerations. On the contrary, by distracting management from their core competencies, shareholder welfarism would ultimately lower shareholder welfare.

The future of CSR, as with its past, is instead with enlightened shareholder value (“ESV”). But the existing law-and-economics literature on ESV has been stunted by key misconceptions, which we attempt to dispel. The increasing use by various actors in the corporate system of normative arguments that sound in ESV terms may lead to new pathways for achieving social progress.

INTRODUCTION

Corporate managers play crucial roles in our society, sitting as they do atop organizations in control of vast agglomerations of resources. A long-standing debate in American law concerns how corporate fiduciaries should conceive of their jobs—what objective should they pursue? The traditional understanding is that the fiduciaries of a business corporation should pursue shareholder value, and much of our corporate governance system is designed to that end. Pursuit of shareholder value, of course, can conflict with other interests in society. The classic alternative to the shareholder value maximization paradigm is some form of stakeholderism, in which shareholder wealth is but one of the ends to be sought by management, alongside the interests of workers, other suppliers, customers, and the broader community.

But stakeholderism has foundered due to two key problems. First, state corporation statutes give shareholders the right to elect the board of directors, which in turn holds legal power to manage the corporation.1See, e.g., Del. Code Ann. tit. 8, §§ 141(a), 211(b) (2023). Directors are naturally oriented toward serving the interests of their equity investor electorate, so that absent deeper reforms that would give other stakeholders board representation, shareholders’ interests are likely to continue to be treated as primary.2Leo E. Strine, Jr., Corporate Power Is Corporate Purpose I: Evidence from My Hometown, 33 Oxford Rev. Econ. Pol’y 176, 179 (2017); Lucian A. Bebchuk & Roberto Tallarita, The Illusory Promise of Stakeholder Governance, 106 Cornell L. Rev. 91, 146 (2020); Edward B. Rock, For Whom Is the Corporation Managed in 2020?: The Debate over Corporate Purpose, 76 Bus. Law. 363, 394 (2021). Second, stakeholder theorists have not congealed around any methodology to determine how corporate management should strike the inevitable trade-offs among the competing interests of different stakeholders, simply leaving it up to management to sort out as they see fit.3Margaret M. Blair & Lynn A. Stout, Director Accountability and the Mediating Role of the Corporate Board, 79 Wash. U. L.Q. 403, 408 (2001). Lacking any metric against which management performance can be judged, stakeholderism in practice risks reducing the accountability of management.4Frank H. Easterbrook & Daniel R. Fischel, The Economic Structure of Corporate Law 38 (1991); Michael C. Jensen, Value Maximization, Stakeholder Theory, and the Corporate Objective Function, 14 J. Applied Corp. Fin., Fall 2001, at 8, 14 (2001) (“By failing to provide a definition of better [and worse decision-making], stakeholder theory effectively leaves managers and directors unaccountable for their stewardship of the firm’s resources.”).

The debate about corporate purpose is old, dating back at least as far as the foundational exchange between E. Merrick Dodd Jr. and A.A. Berle Jr. in the pages of the Harvard Law Review in the early 1930s.5See generally E. Merrick Dodd, Jr., For Whom Are Corporate Managers Trustees?, 45 Harv. L. Rev. 1145 (1932); A.A. Berle, Jr., For Whom Corporate Managers Are Trustees: A Note, 45 Harv. L. Rev. 1365 (1932). Yet as early as that era, there were those who questioned the extent to which shareholder interests are actually incompatible with stakeholder interests. Mistreating workers, customers, and other firm patrons is not in general a recipe for long-term business success.6Jensen, supra note 4, at 16 (“[I]t is a basic principle of enlightened value maximization that we cannot maximize the long-term market value of an organization if we ignore or mistreat any important constituency.” (emphasis omitted)). As Dodd himself put it, “No doubt it is to a large extent true that an attempt by business managers to take into consideration the welfare of employees and consumers . . . will in the long run increase the profits of stockholders.”7Dodd, supra note 5, at 1156. While not embraced by Dodd,8Id. at 1156–57 (“[O]ne need not be unduly credulous to feel that there is more to this talk of social responsibility on the part of corporation managers than merely a more intelligent appreciation of what tends to the ultimate benefit of their stockholders.”). this so-called “enlightened” shareholder value view has historically represented the primary alternative to stakeholderism for those seeking to reorient corporate managers toward more socially responsible business practices.9See Dorothy S. Lund, Enlightened Shareholder Value, Stakeholderism, and the Quest for Managerial Accountability in Research Handbook on Corporate Purpose and Personhood 91, 94–99 (Elizabeth Pollman & Robert B. Thompson eds., 2021) (documenting embrace of ESV among corporate managers and investors).

But recent years have given rise to new perspectives on how corporate managers should understand shareholders’ interests that aim to weaken the grip of shareholder value on the hearts and minds of corporate managers and provide a new north star by which they could chart a more socially responsible course. The key to these innovations is the recognition that the shareholders of a business corporation in general care about more than just the return on the company’s common stock. For one, shareholders care about other stakeholders’ interests directly because of their own personal normative commitments (their “social preferences,” in the reductive parlance of economists). And even from just a financial perspective, each shareholder’s stake in the company is held as part of a broader portfolio. Some portion of the external harms that arise as by-products of the company’s pursuit of profits—to the environment, for example—will ultimately fall on other companies held in shareholders’ portfolios. Under this view, for corporate fiduciaries to further shareholders’ true interests, properly understood, they must eschew narrow shareholder value maximization and instead focus on shareholder welfare maximization, which incorporates these shareholder social preferences and portfolio effects.

In this Article we provide the first comprehensive analysis of these attempts, new and old, to pursue corporate social responsibility through shareholder governance. In Part I, we provide a brief overview of the traditional debate about the objective of a business corporation. In Part II, we dilate on the idea of enlightened shareholder value (“ESV”) as a way to pursue corporate social responsibility (“CSR”) within the traditional norm of shareholder primacy. In Part III, we outline the more recent attempts to improve corporate conduct by incorporating more holistic understandings of shareholder interests, one that focuses on shareholders’ social concerns and another that considers shareholders’ financial interests from a diversified portfolio perspective, which we refer to as the shareholder social preferences (“SSP”) view and the portfolio value maximization (“PVM”) view, respectively.

In Part IV, we turn to evaluating the extent to which these three competing approaches to pursuing CSR through shareholder governance—ESV, SSP, and PVM—are likely to induce public companies to incur costs on a voluntary basis in ways that further the interests of other stakeholders in the firm. We refer to such actions as engaging in CSR. We begin by analyzing the degree to which the corporate objective posited by each approach captures CSR concerns, ignoring the challenges to inducing managers to pursue each objective. While the long-term shareholder value objective of ESV does align to some extent with key stakeholder concerns, it falls short of resolving all social conflicts about corporate conduct, even if we put feasibility concerns to the side. But incorporating shareholders’ social preferences into the corporate objective offers little hope for improvement. For one, shareholder welfare puts far greater relative weight on long-term shareholder value than would a proper conception of social welfare. As well, shareholders’ insulation from the social and moral pressures that generate prosocial behavior at the individual level mutes their social preferences with respect to corporate conduct. Finally, conflicts among shareholders about social issues further dampen the role of social preferences in shareholder welfare.

Diversified shareholders’ portfolio value is even less normatively attractive as a corporate objective. It captures only a small portion of the externalities like pollution that its proponents hope to address. The type of externalities it does capture effectively are competitive effects on other firms—like competitors’ loss of business following a cut to the price of the firm’s output—the result of which is to motivate socially destructive anticompetitive conduct.

We then consider the feasibility of implementing each approach. While ESV is substantially feasible in terms of its information demands, management’s incentives are more mixed due to standard agency problems. Corporate short-termism is one type of agency cost that might result in management failing to engage in CSR that would benefit shareholders in the long-term. Overinvestment due to empire building in high-negative externality industries is another. In sum, in practice management will sometimes, perhaps often, fall short of the degree of social responsibility that is consistent with the shareholder value objective.

Adding shareholders’ social preferences to the corporate objective, however, would provide little by way of incremental incentives to act responsibly. For one, given that shareholders’ social preferences are in important part associative, the shareholders actually willing to hold the shares of the companies that pose the greatest social concerns will be those least concerned about the social issues implicated. As well, management faces significant information problems in gleaning the strength and content of the social preferences of their shareholder base. Indeed, diversified shareholders themselves, we submit, would struggle to formulate such preferences across the myriad social issues implicated by their portfolios. These information problems of the SSP approach in turn produce a fundamental incentive problem. With one far more important component of the objective for which managers have reasonably good information—shareholder value—and one far less important component for which they have little information—shareholders’ social preferences—the optimal incentive scheme focuses management squarely on shareholder value. Attempts to push management to attend to shareholders’ social preferences thus risk doing more harm to shareholder (and social) welfare than good by distracting management from their core competencies.

The story is much the same for PVM. Corporate managers are likely to be far better informed about how their business produces cash flows for the company and about competitive effects on other firms than about other externalities of the company’s business on other companies. Nor are institutional investors likely to be in a meaningfully better position to provide information on portfolio externalities to managers. The optimal incentive scheme for firm managers under PVM would thus also focus on long-term shareholder value of the firm. To the extent it would incorporate externalities, they would be largely of the competitive variety, leading to worse corporate behavior from a social perspective.

To be sure, one might seek to sidestep these managerial incentive and information problems by simply devolving greater corporate control to shareholders, and a number of prominent scholars have indeed advocated taking such a direct approach to implementing shareholder welfarism.10See infra Section IV.C. However, for publicly traded corporations at the center of these proposals, the basic economic logic of centralized management would continue to apply, suggesting any such departure from centralized management would entail sacrificing many of the efficiencies that have long justified this form of corporate organization. As well, recent work in economics suggesting that shareholders would act like social planners were they to have greater voting rights on operational decisions is based on strong assumptions and is in practice implausible. Devolving corporate control to shareholders would therefore offer little benefit in terms of more responsible corporate conduct and would entail substantial costs.

Shareholder governance does hold significant promise for improving corporate conduct, but this promise does not stem from any innovation in our basic understanding of shareholders’ interests along the lines of shareholder welfarism. Rather, the future of CSR, as with its past, is with ESV. The existing law-and-economics literature on ESV, however, has been stunted by two key misconceptions, which we attempt to dispel in Part V. The first is to frame ESV as an alternative to shareholder value as a corporate objective. This is a category mistake: ESV is best understood as a reform agenda targeting a particular class of agency costs that harm not only shareholders but also other corporate stakeholders. A second misconception is that the behavior of all the key actors in the corporate system is fully determined by their incentives and so ideas inspired by ESV cannot improve it. But we show that this determinacy paradox is a challenge for all normative arguments in corporate law scholarship. The generality of this analytic challenge for normative arguments in the field has not previously been recognized. Yet we also provide good reasons to think that this challenge can be surmounted in the case of ESV. We conclude by outlining a research agenda on ESV that would help illuminate the scope for further improvements to CSR through shareholder governance.

I.  THE TRADITIONAL DEBATE ABOUT CORPORATE OBJECTIVE

The traditional debate about the objective of a business corporation traces back to an influential exchange almost a century ago between Columbia Law School Professor Adolf A. Berle and Harvard Law School Professor E. Merrick Dodd that grappled with a fundamental question posed by the publicly traded corporation: Given the practical inability of dispersed shareholders to monitor managers, what maximand should managers pursue in exercising their resulting wide discretion over corporate affairs?11See Dodd, supra note 5, at 1147 (“Directors and managers of modern large corporations . . . are free from any substantial supervision by stockholders by reason of the difficulty which the modern stockholder has in discovering what is going on and taking effective measures even if he has discovered it.”).

A.  Shareholder Wealth Maximization

Berle’s solution was to turn to the law of trusts and argue that managers are trustees obligated to exercise their discretion solely for the benefit of the shareholders,12See A.A. Berle, Jr., Corporate Powers as Powers in Trust, 44 Harv. L. Rev. 1049, 1049 (1931). which he understood narrowly in terms of their interests in the corporation’s profits.13Berle, supra note 5, at 1367 (“Now I submit that you can not abandon emphasis on ‘the view that business corporations exist for the sole purpose of making profits for their stockholders’ until such time as you are prepared to offer a clear and reasonably enforceable scheme of responsibilities to someone else.”). It was this view of the corporation that was later reprised in Milton Friedman’s famous assertion that corporate executives’ “responsibility is to conduct the business in accordance with [shareholders’] desires, which generally will be to make as much money as possible while conforming to the basic rules of the society.”14Milton Friedman, A Friedman Doctrine—The Social Responsibility of Business Is To Increase Its Profits, N.Y. Times (Sept. 13, 1970), https://www.nytimes.com/1970/09/13/archives/a-friedman-doctrine-the-social-responsibility-of-business-is-to.html [https://perma.cc/NSE6-ZBZU]. For Berle, this was a matter of managerial accountability. The only alternative he saw to the shareholder wealth maximization norm was to simply hand over “the economic power now mobilized and massed under the corporate form . . . to the present administrators with a pious wish that something nice will come out of it all.”15Berle, supra note 5, at 1368.

The shareholder wealth maximization norm has historically enjoyed broad support for several reasons. First, as a matter of economic theory, if markets are complete, firms are price takers, and there are no externalities not effectively addressed by government policy, corporate profit maximization results in a socially efficient outcome in the sense that there is no way to improve anyone’s well-being without making someone else worse off.16See Kenneth J. Arrow & Gerard Debreu, Existence of an Equilibrium for a Competitive Economy, 22 Econometrica 265, 265 (1954). By running the firm to maximize the value of the residual claims, the social pie is also maximized so long as government policy addresses externalities. Under the traditional shareholder value maximization view, then, externalities and distributive concerns are appropriately addressed by government policy, not by corporate managers assuming responsibility for them. Similarly, under these conditions, shareholders with conflicting preferences about the timing of consumption will nevertheless be unified in a corporate mandate to maximize shareholder wealth, since shareholders can satisfy their diverse consumption preferences by borrowing and saving.17See generally Steinar Ekern & Robert Wilson, On the Theory of the Firm in an Economy with Incomplete Markets, 5 Bell J. Econ. & Mgmt. Sci. 171 (1974)(explaining that with complete markets for borrowing and saving, it is in the interest of each shareholder to maximize firm value). Second, these theoretical arguments are complemented by the agency cost concerns articulated by Berle. Share value provides a simple metric by which to evaluate managers and to hold them accountable for the efficient deployment of corporate assets. Indeed, pioneering work on agency cost theory by Michael Jensen and William Meckling in the 1970s later formalized Berle’s central premise.18See Michael C. Jensen & William H. Meckling, Theory of the Firm: Managerial Behavior, Agency Costs and Ownership Structure, 3 J. Fin. Econ. 305, 312 (1976). Lastly, the basic structure of corporate law reflects the shareholder value maximization norm, particularly in the key state of Delaware. While legal authority to manage the corporation is lodged in its board of directors, it is the stockholders who are entitled to elect directors.19See, e.g., Del. Code Ann. tit. 8, § 141(a) (2023); Model Bus. Corp. Act § 8.01(b) (2023) (establishing that business and affairs of corporations shall be managed by or under direction of board of directors); Del. Code Ann. tit. 8, § 211(b) (2023) (“[A]n annual meeting of stockholders shall be held for the election of directors on a date and at a time designated by or in the manner provided in the bylaws.”). Likewise, courts have defined the fiduciary duties that directors owe to the corporation as ultimately oriented toward stockholder wealth.20As summarized by Vice Chancellor Laster in In re Trados, Inc., “the standard of conduct for directors requires that they strive in good faith and on an informed basis to maximize the value of the corporation for the benefit of its residual claimants [that is, common stockholders] . . . not for the benefit of its contractual claimants.” In re Trados, Inc., 73 A.3d 17, 40–41 (Del. Ch. 2013). A broad range of complementary institutions has developed that further entrench shareholder interests as the primary end of the corporate system.21Dorothy S. Lund & Elizabeth Pollman, The Corporate Governance Machine, 121 Colum. L. Rev. 2563, 2575–78 (2021).

B.  Stakeholderism

In contrast to Berle, Dodd identified a trend in public opinion toward viewing the publicly held corporation as an “economic institution which has a social service as well as a profit-making function”22Dodd, supra note 5, at 1148. and believing that “business has responsibilities to the community.”23Id. at 1153. He viewed this trend in public opinion as desirable and likely to become the view of corporate managers, who would develop business ethics that would be “in some degree those of a profession rather than of a trade.”24Id. at 1161. Normatively he argued against the position of Berle that corporate fiduciaries have a legal responsibility just to stockholders in order to preserve the freedom of action necessary for management to fulfill their inchoate social obligations.25Id. The conceptualization of those to whom corporate managers owe these social responsibilities as stakeholders took off much later with an influential book aimed at corporate managers by Edward Freeman titled Strategic Management: A Stakeholder Approach.26R. Edward Freeman, Strategic Management: A Stakeholder Approach (1984). Freeman offered a capacious definition of stakeholders as “any group or individual who can affect or is affected by the achievement of the organization’s objectives.”27Id. at 46. Owing in part to the influence of Freeman,28Joshua D. Margolis & James P. Walsh, Misery Loves Companies: Rethinking Social Initiatives by Business, 48 Admin. Sci. Q. 268, 279 (2003) (“Freeman’s ideas provided a language and framework for examining how a firm relates to ‘any group or individual who can affect or is affected by the achievement of the organization’s objective.’ ”). the school of thought originally launched by Dodd has since become known as “stakeholder theory” or simply “stakeholderism.”29Bebchuk & Tallarita, supra note 2, at 94. Under this view, corporate fiduciaries should voluntarily advance not just the interests of shareholders but also the interests of workers, creditors, other suppliers, customers, and all others who are affected by the corporation’s activities. The term “corporate social responsibility” is generally used to refer to this view of a firm’s obligations to advance the interests of its stakeholders.

To organize the various types of social concerns that animate stakeholder theory, it is useful to distinguish between corporate stakeholders that transact with the firm—which we will refer to as firm patrons—and stakeholders that do not. One type of concern regarding the treatment of firm patrons stems from market failures that lead to inefficient outcomes. A primary source of such market failures is market power. A firm with market power in the labor market, for example, will depress workers’ wages in order to maximize its profits.30Efraim Benmelech, Nittai K. Bergman & Hyunseob Kim, Strong Employers and Weak Employees: How Does Employer Concentration Affect Wages?, 57 J. Hum. Res. S200, S201 (2022). Similarly, market power with respect to its customers can lead to inefficiently high prices for the firm’s output.31Robert S. Pindyck & Daniel L. Rubinfeld, Microeconomics 359 (6th ed. 2005). In both cases these deviations from competitive prices result in deadweight costs—inefficient reductions in transactions in the market. Market power also raises distributive concerns—a greater share of the social surplus generated in the relevant market goes to the firm rather than firm patrons. Distributive concerns can also arise even in the absence of market power when the relevant market is competitive and efficient. Stakeholderists might view the low wages in a competitive labor market, for example, as socially undesirable and advocate for the firm to pay its workers more.32See, e.g., Addie Stone, Improving Labor Relations Through Corporate Social Responsibility – Lessons from Germany and France, 46 Cal. W. Int’l L.J. 147, 150–51 (2016) (“Employees are key stakeholders, and their compensation is an important CSR issue. . . . [C]ompanies should focus their CSR efforts on providing a living wage to its employees.”).

Concerns about non-firm patrons, in contrast, typically involve externalities. Consider, for example, climate change. Firms’ operations inevitably entail some amount of greenhouse gas emissions, which contribute to the total stock of greenhouse gases in the atmosphere and in turn to the warming of the planet. The global scope of the climate change problem, in terms of both its causes and effects, means that essentially the entire global community is affected by every firm’s operations and hence can be considered a stakeholder of every firm. But many other externalities are much smaller in scale, resulting in a firm’s local community typically having a greater interest in the firm’s operations than those further afield.

Note that the basic normative claim at the heart of stakeholderism—that corporate fiduciaries should voluntarily advance the interests of all firm stakeholders and not just the interests of shareholders—presumes some sort of imperfection in current law and policy or in corporations’ responses to it. Stakeholderists argue, in effect, that current public policy is not sufficient to protect stakeholder interests, and so corporate managers should go even further on their own.33See David L. Engel, An Approach to Corporate Social Responsibility, 32 Stan. L. Rev. 1, 36 (1979) (“One cannot persuasively claim to have found an extra-profit goal that society wants corporations to pursue, unless one can offer at least a plausible explanation of why the legislature did not long ago enact liability rules, regulations, or other measures, to implement the goal in question quite independently of any management practice of social responsibility.”).

Notwithstanding the orientation of corporate law toward shareholder wealth maximization, certain core features of corporate law provide the managerial discretion that is necessary to implement stakeholderism. Director decision-making in the absence of financial conflicts of interest remains largely shielded from judicial scrutiny by the business judgment rule. As a result, corporate managers enjoy broad discretion to consider an array of stakeholder interests so long as their decisions can be justified as ostensibly in the interests of the corporation.34See, e.g., Shlensky v. Wrigley, 237 N.E.2d 776, 780 (Ill. App. Ct. 1968) (holding that, absent fraud, illegality, or conflict of interest, the decision of the Chicago Cubs not to hold night games was properly in the hands of the board of directors and the courts would not intervene). The court pointed out that the decision might in principle be justified based on the financial interests of the corporation, for example, because of the possible negative effect on the property value of Wrigley Field that a deterioration in the surrounding neighborhood might cause. Id. Moreover, many state legislatures have amended corporate statutes to increase the compatibility of corporate law with stakeholderism. For instance, so-called constituency statutes have been adopted in most states—but not Delaware—that make clear that corporate fiduciaries are not required to consider only shareholder interests to the exclusion of other stakeholders’ interests.35Margaret M. Blair, Ownership and Control: Rethinking Corporate Governance for the Twenty-First Century 218–19 (1995). The main motivation for these reforms was to prevent corporate takeovers on the ground that takeovers and their associated restructurings could be harmful to workers and local communities.36See, e.g., Eric W. Orts, Beyond Shareholders: Interpreting Corporate Constituency Statutes, 61 Geo. Wash. L. Rev. 14, 23–24 (1992). Even in Delaware, the case law evolved to endorse the prerogative of corporate directors to take action to fend off a premium acquisition offer that the shareholders are eager to accept in order to pursue directors’ long-term vision of what is in the corporation’s best interest.37See Paramount Commc’ns, Inc. v. Time Inc., 571 A.2d 1140, 1142 (Del. 1989) (upholding defensive measures by the Time, Inc. board motivated in part by a desire to preserve the company’s editorial integrity). More recently, the adoption of public benefit corporation statutes has been similarly grounded in a desire to enable business corporations to pursue stakeholderist objectives.38See Jill E. Fisch & Steven Davidoff Solomon, The “Value” of a Public Benefit Corporation, in Research Handbook on Corporate Purpose and Personhood 68, 68 (Elizabeth Pollman & Robert B. Thompson eds., 2021). These developments show that there is nothing inevitable about privileging the interests of investors in operating a commercial enterprise. Indeed, a wide variety of enterprises—such as consumer cooperatives, producer cooperatives, and nonprofits—have chosen to privilege a different set of stakeholders.39Cf. Henry Hansmann, The Ownership of Enterprise (1996) (developing an efficiency-based theory for the assignment of ownership rights to different classes of firm patrons).

II.  ENLIGHTENED SHAREHOLDER VALUE

Stakeholderism correctly identifies that shareholders’ interests in corporate profits can conflict with other interests in society. From a static, short run perspective especially, these conflicts can loom large. Squeezing suppliers and customers can increase corporate profits at their expense. Cutting back on greenhouse gas emissions will improve the environment but at a direct cost to the company’s bottom line. And so on and so forth—the list of such conflicts is endless. But taking a longer-term perspective on the company and its business may lessen the degree of conflict between stockholders and other firm stakeholders. More generally, for a range of reasons, considered in some detail below, it can be in shareholders’ interests for the company to incur costs to improve the well-being of the firm’s stakeholders. Or put more colloquially, companies can “do well by doing good.” This enlightened shareholder value perspective, while often dismissed by stakeholder theorists as insufficient40See, e.g., Dodd, supra note 5, at 1156–57; Colin P. Mayer, Prosperity: Better Business Makes the Greater Good 6–7 (2018) (“ ‘Doing well by doing good’ is a dangerous concept because it suggests that philanthropy is only valuable where it is profitable, and it converts charity into profit-generating entities . . . .”). and by shareholder value theorists as uninteresting41See, e.g., Einer Elhauge, Sacrificing Corporate Profits in the Public Interest, 80 N.Y.U. L. Rev. 733, 744 (2005); Bebchuk & Tallarita, supra note 2, at 110 (“Enlightened shareholder value is thus no different from shareholder value tout court.”). or even counterproductive,42Lucian A. Bebchuk, Kobi Kastiel & Roberto Tallarita, Does Enlightened Shareholder Value Add Value?, 77 Bus. Law. 731, 734 (2022). has gained increasing traction in recent years as a way to respond to the concerns of stakeholderism that is compatible with existing institutions that put shareholder interests first.43See, e.g., Lund, supra note 9, at 97–98 (arguing that concerns about corporate short-termism have led to a shift toward an enlightened shareholder value perspective); Jensen, supra note 4, at 9 (“Enlightened value maximization uses much of the structure of stakeholder theory but accepts maximization of the long-run value of the firm as the criterion for making the requisite tradeoffs among its stakeholders . . . . In so doing, it solves the problems arising from the multiple objectives that accompany traditional stakeholder theory by giving managers a clear way to think about and make the tradeoffs among corporate stakeholders.”); Michael E. Porter & Mark R. Kramer, Creating Shared Value, 89 Harv. Bus. Rev., Jan.–Feb. 2011, at 62, 64–65; Alex Edmans, Grow the Pie: How Great Companies Deliver Both Purpose and Profit 55–56 (2020).

Today the idea of ESV is more commonly referred to under the moniker “ESG,” which stands for “Environmental, Social, and Governance.”44See The Global Compact, Who Cares Wins, at 3 (2004). While ESG is a notoriously protean term, used for a range of different ideas,45For an illuminating discussion of the origins of and diverse meanings ascribed to ESG, see generally Elizabeth Pollman, The Making and Meaning of ESG, Harv. Bus. L. Rev. (forthcoming), https://papers.ssrn.com/abstract=4219857 [https://perma.cc/3JCD-LP55]. its origins are as a term that captures ways that investors can improve their risk-adjusted returns by incorporating environmental, social, and governance considerations into their investment process.46See id. at 11–13; The Global Compact, supra note 44, at i–ii (2004); Alex Edmans, The End of ESG, 52 Fin. Mgmt. 3 (2022). A key aspect of the standard rationale for the use of ESG factors to improve investment returns is the idea that such factors affect profitability at the level of the portfolio company.47The Global Compact, supra note 44, at 9; Robert G. Eccles, Ioannis Ioannou & George Serafeim, The Impact of Corporate Sustainability on Organizational Processes and Performance, 60 Mgmt. Sci. 2835, 2849, 2851 (2014) (finding high sustainability companies outperform low sustainability companies both in terms of stock market and accounting performance). Indeed, the notion that paying attention to ESG matters for firm financial performance has become part of the zeitgeist of recent years, with public companies increasingly discussing their ESG initiatives on quarterly earnings calls,48Goldman Sachs Equity Research, The Corporate Commotion – A Rising Presence of ESG in Earnings Calls 25 (2020), https://www.goldmansachs.com/insights/pages/gs-sustain-corporate-commotion-f/report.pdf [https://perma.cc/AUE3-3YTN]. hiring executives to oversee ESG reforms,49See Stavros Gadinis & Amelia Miazad, Corporate Law and Social Risk, 73 Vand. L. Rev. 1401, 1420 (2020). and tying executive compensation to ESG metrics.50The Conference Board, Linking Executive Compensation to ESG Performance 3 (2022), https://www.conference-board.org/pdfdownload.cfm?masterProductID=41301 [https://perma.
cc/Z2M6-7NCV] (reporting that 73% of S&P 500 companies tied executive compensation to some form of ESG performance as of 2021).
Another aspect of this rationale for ESG investing is the claim that the stock market misprices ESG factors.51See Max M. Schanzenbach & Robert H. Sitkoff, Reconciling Fiduciary Duty and Social Conscience: The Law and Economics of ESG Investing by a Trustee, 72 Stan. L. Rev. 381, 437 (2020) (“For an investor to be able to profit by trading on ESG factors, the market must consistently misprice them.”). To be sure, the term ESG is also used for practices that sacrifice investor returns in order to achieve benefits for stakeholders.52See id. at 397–98 (referring to this form of ESG as “collateral benefits ESG”). But in the main, much of the standard rhetoric around ESG, and its intellectual origins, reflect what we refer to as ESV.53See, e.g., United Nations Principles for Responsible Investment, A Blueprint for Responsible Investment 7 (2017), https://www.unpri.org/download?ac=5330 [https://perma.cc/
W4L7-9FPP] (“That environmental, social and governance factors each contribute to creating long-term value is a case well-understood by many, but remains new to many others – so it is a case we must continue to make.”).
As of 2022, some $8.4 trillion in assets under management in the United States are invested using an ESG approach.54US SIF Foundation, 2022 Report on US Sustainable Investing Trends 2 (2022).

ESV theorists typically describe the corporate objective as long-term shareholder value. The modifier long-term serves two purposes. First, it signifies that much of the financial value of the firm’s shares stems from cash flows it will produce well into the future. Second, it reflects the possibility that a company’s stock price might not fully reflect immediately the future cash flows that an action to sacrifice corporate cash flows today will ultimately produce.55See Jensen, supra note 4, at 17; Edmans, supra note 43, at 121. But the basic valuation framework underlying ESV is entirely conventional: the firm should be managed to maximize the net present value of the firm’s equity, calculated by discounting the cash flows available to equity holders using the appropriate risk-adjusted discount rate (however long it might take for the markets to catch up and price the company’s stock accordingly). In other words, ESV is not an alternative conception of corporate purpose—it retains the exact same corporate objective as standard shareholder value theory.56Analyses of ESV as a distinct normative standard for corporate decision-making thus largely miss the point of ESV. See generally, e.g., Bebchuk et al., supra note 42. We discuss critiques of ESV in some detail in Part V infra. Instead, ESV theory identifies a set of mechanisms through which firm managers can increase long-term shareholder value by behaving in a more socially responsible way.57For instance, a recent McKinsey Quarterly publication identifies five distinct channels through which more socially responsible corporate behavior can improve long-term profitability. Witold Henisz, Tim Koller & Robin Nuttall, Five Ways that ESG Creates Value, McKinsey Q., Nov. 2019, at 4.

With respect to the treatment of firm patrons, one mechanism posited entails a type of efficiency wage: treating a class of firm patrons better can induce reciprocal improved treatment of the firm by those firm patrons. For example, when a firm pays its workers better than their outside option—the market wage for similar labor—workers have greater incentive to perform their jobs well, in order to reduce the risk of dismissal, and the resulting increase in productivity can more than compensate for the firm’s increased wage bill.58See Carl Shapiro & Joseph E. Stiglitz, Equilibrium Unemployment as a Worker Discipline Device, 74 Am. Econ. Rev. 433, 433–34 (1984). Other accounts emphasize the importance of employee morale and perceptions of fairness: workers who are paid what they consider to be an unfair wage are likely to shirk or otherwise cut back on effort and vice versa.59George A. Akerlof & Janet L. Yellen, The Fair Wage-Effort Hypothesis and Unemployment, 105 Q. J. Econ. 255, 263 (1990). Similarly, a corporation that invests in promoting a diverse and inclusive work culture might boost employee motivation and performance60See Deloitte, Waiter, Is that Inclusion in My Soup?: A New Recipe To Improve Business Performance 4 (2013), https://www2.deloitte.com/content/dam/Deloitte/au/Documents/
human-capital/deloitte-au-hc-diversity-inclusion-soup-0513.pdf [https://perma.cc/5D8G-PHNA]; Jie Chen, Woon Sau Leung & Kevin P. Evans, Female Board Representation, Corporate Innovation and Firm Performance, 48 J. Empirical Fin. 236, 237 (2018).
and attract talented workers away from less enlightened competitors.61Gail Robinson & Kathleen Dechant, Building a Business Case for Diversity, 11 Acad. Mgmt. Exec. 21, 25 (1997). Consistent with this view—and with the stock market underpricing the benefits of favorable treatment of workers—the shares of companies identified as among the “100 Best Companies to Work For in America” earned significant excess returns from 1994 to 2009.62Alex Edmans, Does the Stock Market Fully Value Intangibles? Employee Satisfaction and Equity Prices, 101 J. Fin. Econ. 621, 621 (2011) [hereinafter Edmans, Does the Stock Market Fully Value Intangibles?]; see Alex Edmans, The Link Between Job Satisfaction and Firm Value, with Implications for Corporate Social Responsibility, 26 Acad. Mgmt. Persp. 1, 11 (2012) [hereinafter Edmans, The Link Between Job Satisfaction and Firm Value].

A related mechanism stems from the value of inducing firm-specific investments from firm patrons. A firm’s contracts with its patrons are often long-term and, in important respects, implicit.63See Oliver E. Williamson, The Economic Institutions of Capitalism: Firms, Markets, Relational Contracting 194 (1985). Workers, for example, invest in human capital that is to some extent specific to the firm and less valuable elsewhere. In order to induce workers to make such costly investments, the firm promises in return to pay them a share of the surplus generated by their increased productivity. For such relational contracts to work, however, firm patrons must be able to trust the firm to perform its end of the bargain down the line. Breaching that implicit contract by cutting wages, say, can ultimately harm shareholders by destroying the firm’s reputation for trustworthiness.64Andrei Shleifer & Lawrence H. Summers, Breach of Trust in Hostile Takeovers, in Corporate Takeovers: Causes and Consequences 33, 37–38 (Alan J. Auerbach ed., 1988). Implicit contracts and the value of the firm’s reputation can also provide reasons for the firm to act in a socially responsible manner with respect to its customers. Consider a car insurance company that can increase its profits in the short run by engaging in various practices that slow down or limit the payment on policyholders’ claims. Such short-term financial benefits, however, might be swamped by the future costs of lost customers from the resulting harm to the firm’s reputation as a reliable insurer that treats its policyholders fairly.

The ESV perspective also posits a set of mechanisms through which incurring costs to treat non-patrons well can ultimately create net financial benefits to shareholders. Consider, for example, an energy company’s decision of how much to invest in exploring for oil. The optimal level of investment if one takes a myopic view and assumes that the current market demand for oil will continue indefinitely might be much higher than if one instead adopts a more realistic forecast of the coming transition to a low-carbon economy due to future policy changes and technological developments. The idea is that putting one’s head in the ground and investing based on a naïve assumption of continuing demand, even if it generates increased profits in the short- to medium-term, risks the eventual incurrence of large losses on stranded assets.

The social preferences of one class of firm patrons can also produce financial incentives to treat other classes of firm patrons and non-patrons well.65The social views of Millennial and Gen Z workers and customers might produce greater incentive for firms to engage in more socially responsible behavior than in the past, given their evidently greater willingness to express those views in their decisions about where to work and shop. See Michal Barzuza, Quinn Curtis & David H. Webber, The Millennial Corporation, 28 Stan. J.L. Bus. & Fin. 255, 259–61 (2023). For instance, given consumer demand for environmentally sustainable products, investment in these products can result in increased profits as well as an improved environment.66Stephanie M. Tully & Russell S. Winer, The Role of the Beneficiary in Willingness to Pay for Socially Responsible Products: A Meta-Analysis, 90 J. Retailing 255, 265 (2014).

While the foregoing identifies conceptually coherent mechanisms through which incurring costs to further stakeholder interests can ultimately redound to the financial benefit of stockholders, we do not mean to suggest that all corporate decisions ostensibly justified on that basis are in fact in stockholder interests. Indeed, ESV arguments might be advanced strategically by stakeholderists for actions that in fact will reduce long-term shareholder value. Similarly, ESV might be used as cover by management for actions taken to further management’s interests at the expense of stockholders.67Jonathan Macey, Why Is the ESG Focus on Private Companies, Not the Government?, Bloomberg L. (Aug. 19, 2021, 1:01 AM), https://news.bloomberglaw.com/esg/why-is-the-esg-focus-on-private-companies-not-the-government [https://perma.cc/C8ZM-4Y3Q] (“Managers like ESG investing because the concept is so complex and multi-faceted that almost any action short of theft or outright destruction of corporate property can be defended on some ESG ground or the other.”). We return to the information and incentive problems posed by ESV in Part IV below.

III.  SHAREHOLDER WELFARISM

The ESV view posits considerable alignment between the financial interests of shareholders in the long-term and the interests of other firm patrons and the broader society. It thus provides one avenue to pursue CSR through shareholder governance. We now consider an alternative approach to doing so that is newer to the scene, which we refer to as shareholder welfarism. It posits that corporate management should seek to maximize shareholder welfare, not just share value, by incorporating a more complete understanding of how the corporation affects the well-being of shareholders. There are two primary strands of shareholder welfarism in the literature—the shareholder social preferences view and the portfolio value maximization view—which we take up in turn.68A third version of what we call shareholder welfarism focuses on the direct effects of corporate externalities on the well-being of shareholders—for example, shareholders’ health may be harmed by corporate pollution. See Michael Simkovic, Natural-Person Shareholder Voting, 109 Cornell L. Rev. (forthcoming 2024) (manuscript at 4), https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4180982# [https://perma.cc/H52C-ZAV3]. We view this as a less significant component of shareholder welfare, in part because the wealth generated through share ownership may enable shareholders to avoid exposure to many corporate externalities. Id. Moreover, much of our analysis of SSP and PVM apply to the direct effects component as well, so we omit treatment of this version in the interest of brevity.

A.  Shareholder Social Preferences

The shareholder social preferences (“SSP”) version of shareholder welfarism begins with the commonsense observation that public company shareholders care about more than just their own wealth—they also have ethical and social concerns. Many shareholders care about the environment, inequality, and racial justice, to give just a few examples, based on their own personal normative commitments. There is of course a wide range of views on such social issues. But while public company shareholders might not be perfectly representative of the entire population, there is no reason to think that corporate shareholders, unlike others in society, are narrowly self-interested and lack any social preferences.

Many shareholders would thus presumably often prefer that company management sacrifice share value in order to further their social preferences, at least to some extent. Consumer markets provide a useful analogy. Consider fair trade coffee, which is sold in major grocery chains across the United States. Fair trade goods are marketed to consumers at a premium price on the basis that the greater markup is passed on to poor producers. This is intended to appeal to consumers with ethical concerns about the treatment of such producers. Such a consumer might be willing to pay more for goods that promise better outcomes for the producers, a hypothesis confirmed by experimental evidence.69The leading study found that replacing a generic product label with a Fair Trade label increases sales of coffee by almost 10%, with higher demand holding steady at up to an 8% price premium. Jens Hainmueller, Michael J. Hiscox & Sandra Sequeira, Consumer Demand for Fair Trade: Evidence from a Multistore Field Experiment, 97 Rev. Econ. & Stat. 242, 253 (2015). Suppose those same consumers are also shareholders of a corporation that sources coffee beans. The SSP view posits that those same social preferences would also lead them to be willing to sacrifice investment returns as shareholders in order for the corporation to pay producers more.70There is some evidence, however, that individuals are less willing to pay to advance social concerns in investment decisions than in consumption decisions. See Scott Hirst, Kobi Kastiel & Tamar Kricheli‐Katz, How Much Do Investors Care About Social Responsibility?, 2023 Wis. L. Rev. 977,  1011. Under the SSP view, corporate fiduciaries should manage the corporation not to maximize shareholder wealth but rather to maximize shareholder welfare, incorporating shareholders’ social preferences.71Oliver Hart & Luigi Zingales, Companies Should Maximize Shareholder Welfare Not Market Value, 2 J.L. Fin. & Acct. 247, 263 (2017).

To be sure, in some cases, shareholder welfare so conceived is in fact maximized by simply maximizing shareholder wealth. Corporate charitable contributions provide an example. Tax complications aside, the goal of furthering shareholder social preferences provides no basis for such corporate philanthropy since the corporation could instead pay those funds out to shareholders, who in turn could donate directly to charity. Oliver Hart and Luigi Zingales—prominent proponents of the SSP view—characterize this as a case in which the social concern is “separable” from the company’s business.72Id. at 249. But Hart and Zingales argue convincingly that social concerns and moneymaking by the company are often inseparable.73Id. They offer as an example shareholder concerns about mass shootings. Walmart might much more effectively advance those shareholder social preferences by no longer selling high-capacity magazines than by contributing the profits from doing so to charity.74Id. Indeed, it seems plausible that for virtually every major CSR concern there are important aspects of the problem that are not completely separable from the businesses of the corporations involved.

The extent to which shareholders are willing to sacrifice their wealth to address various social concerns of course varies from shareholder to shareholder. Hart and Zingales propose that such heterogeneity be handled through voting by shareholders.75Id. at 260–61. The board of directors of the corporation could be required to periodically poll shareholders about corporate policies that implicate social concerns so that the median shareholder’s views on the issue (on a share-weighted basis) prevail. Implicit in this voting-based approach is that the “shareholder welfare” objective weights each shareholder’s preferences by the number of shares they own.76It is not entirely clear how companies with multiple classes of stock with different voting rights and cash flow rights should be handled under the SSP view. One natural approach would be to calculate shareholder welfare by weighting each shareholder’s preferences by the cash flow rights they hold. This would align most closely with the approach taken under the traditional shareholder value view of the corporate objective.

A further wrinkle is that most corporate shares today are held by institutional investors.77Amil Dasgupta, Vyacheslav Fos & Zacharias Sautner, Institutional Investors and Corporate Governance, Founds. & Trends Fin. (forthcoming) (manuscript at 4), https://papers.ssrn.com/sol3/
papers.cfm?abstract_id=3682800 [https://perma.cc/3KG8-MW3Q].
Under the SSP view, it is the social preferences of the underlying investors in those institutions that corporate management should seek to advance. Institutional investors would thus have to channel their investors’ views in voting the stock in their portfolio companies in order for corporate voting to accurately reflect shareholder welfare. Hart and Zingales envision asset managers segmenting the market based on the social views the asset manager will seek to advance in voting shares of its portfolio companies, so that individual investors can simply sort themselves to the appropriate asset manager.78Hart & Zingales, supra note 71, at 265–66. One might wonder whether SSP and shareholder wealth maximization might yield similar results with regard to CSR given the valuation effects of shareholders’ buying and selling stocks according to their social preferences. For instance, if shareholders divest from a dirty company based on their social preferences, the resulting decrease in the company’s stock price might arguably induce wealth-minded managers to turn clean in the name of maximizing shareholder wealth. See Robert Heinkel, Alan Kraus & Josef Zechner, The Effect of Green Investment on Corporate Behavior, 36 J. Fin. & Quantitative Analysis 431, 432–33 (2001). Eleonora Broccardo, Hart, and Zingales argue against this result given that any fall in prices among dirty firms is likely to be muted by marginal investors who purchase the newly discounted shares on account of the lower weight these investors place on their social preferences. See Eleonora Broccardo, Oliver Hart & Luigi Zingales, Exit Versus Voice, 130 J. Pol. Econ. 3101, 3117–20 (2022). Empirical evidence also suggests that divestment from dirty companies produces only modest price declines. See Jonathan B. Berk & Jules H. van Binsbergen, The Impact of Impact Investing 2–3 (L. & Econ. Ctr. at George Mason Univ. Scalia L. Sch., Research Paper No. 22-008, 2021), https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3909166 [https://perma.cc/VJ9N-5Q56]. We discuss sorting of shareholders into firms according to their social preferences infra Section IV.B.2.i.

B.  Portfolio Value Maximization

The portfolio value maximization (“PVM”) strand of shareholder welfarism, in contrast, retains the focus on shareholders’ financial interests from the traditional shareholder value approach but considers their financial interests from a portfolio perspective. Most shareholders in public companies are highly diversified and increasingly so with the ongoing shift from active management to passive investment vehicles.79Vladyslav Sushko & Grant Turner, The Implications of Passive Investing for Securities Markets, BIS Q. Rev., Mar. 2018, at 113, 115. From this perspective, the actual interests of a firm’s shareholders lie in the value of their diversified portfolios, not just in the value of the firm’s shares. Accordingly, corporate fiduciaries should seek to maximize the value of the firm’s shareholders’ portfolios, not their own firm value.

The main implication of the PVM approach concerns between-firm externalities, meaning ways that the decisions of one firm affect the value of other firms. Such spillover effects come in a variety of forms. One form stems from market competition. When a firm gains market share by cutting prices, competing firms often lose customers. Economists refer to this type of external effect as a “pecuniary externality.”80J.-J. Laffont, Externalities, in Allocation, Information and Markets 112, 113 (John Eatwell, Murray Milgate & Peter Newman eds., 1989). A quite different form—referred to as a “technological externality”—occurs when a production or consumption activity imposes costs or benefits on other producers or consumers and does not operate through the price system.81Id. at 112. For example, suppose a factory releases toxic chemicals that reduce agricultural productivity in the surrounding area. From the traditional shareholder value perspective, corporate managers should manage the corporation to maximize the value of its equity without regard to such spillover effects on the value of other firms or on consumers. But under the PVM view, the company’s shareholders would want firm managers to incorporate such external effects to the extent that they reduce the value of other securities held in shareholders’ portfolios.

The social desirability of such PVM behavior by firm managers depends critically on the nature of the externality at issue and the extent to which it is internalized in shareholders’ portfolios. In the case of pecuniary externalities, having firm managers take them into account would interfere with market competition. For example, if each firm in an industry were operated to maximize the total value of the industry, that would entail pricing their output above the competitive level, with all of the standard inefficiencies from monopoly pricing that would result. In recent years a burgeoning empirical literature claims that the growth of diversified institutional investors has in fact led to such anticompetitive outcomes in certain industries.82José Azar, Martin C. Schmalz & Isabel Tecu, Anticompetitive Effects of Common Ownership, 73 J. Fin. 1513, 1558–59 (2018). But a number of papers have raised methodological concerns with this finding. See, e.g., Patrick Dennis, Kristopher Gerardi & Carola Schenone, Common Ownership Does Not Have Anticompetitive Effects in the Airline Industry, 77 J. Fin. 2765, 2766 (2022); Andrew Koch, Marios Panayides & Shawn Thomas, Common Ownership and Competition in Product Markets, 139 J. Fin. Econ. 109, 111 (2021); Katharina Lewellen & Michelle Lowry, Does Common Ownership Really Increase Firm Coordination?, 141 J. Fin. Econ. 322, 324 n.7 (2021). The internalization of pecuniary externalities through the PVM approach is thus generally not socially desirable.

But for technological externalities, PVM offers hope that running the firm in the true interests of shareholders—maximizing the value of their diversified portfolios—would result in more socially responsible corporate behavior. For example, the portfolio value maximizing level of pollution emitted by a firm would take into account the portion of the costs of that pollution that fall on other firms in the portfolio.

These basic implications of running a corporation to maximize the value of diversified shareholders’ portfolios were worked out theoretically by economists decades ago.83See, e.g., Julio J. Rotemberg, Financial Transaction Costs and Industrial Performance 1–3 (Mass. Inst. of Tech. Alfred P. Sloan Sch. of Mgmt., Working Paper No. 1554-84, 1984), https://dspace.mit.edu/bitstream/handle/1721.1/47993/financialtransac00rote.pdf [https://perma.cc/4D
MX-CTC8]; Roger H. Gordon, Do Publicly Traded Corporations Act in the Public Interest? 21–22 (Nat’l Bureau of Econ. Rsch., Working Paper No. 3303, 1990), https://www.nber.org/papers/w3303 [https://perma.cc/KF5Q-VY57]; Robert G. Hansen & John R. Lott, Jr., Externalities and Corporate Objectives in a World with Diversified Shareholder/Consumers, 31 J. Fin. & Quantitative Analysis 43, 44 (1996).
They entered the legal literature when the growth of private and public pension funds, and their growing use of indexed investment strategies, led to calls for these so-called universal owners to exercise their shareholder rights in order to advance broader social interests with respect to corporate behavior.84See James P. Hawley & Andrew T. Williams, The Rise of Fiduciary Capitalism 1–29 (2000); Jeffrey N. Gordon, Systematic Stewardship, 47 J. Corp. L. 627, 632–33 (2022). See generally Robert A.G. Monks & Nell Minow, Watching the Watchers : Corporate Governance for the 21st Century (1996). More recently, Madison Condon has argued that attempts by asset managers to pressure their portfolio companies to combat climate change can be explained by their desire to maximize the value of the diversified portfolios they manage.85Madison Condon, Externalities and the Common Owner, 95 Wash. L. Rev. 1, 26–27 (2020).

The PVM literature has thus largely focused on arguments about how diversified institutional investors should or do exercise their ownership rights in order to change a portfolio company’s policies in ways that increase the value of their diversified portfolios even at the cost of the particular company’s own value.86Id. at 19–26; Gordon, supra note 84, at 658–66. But as Marcel Kahan and Edward Rock argue, responding to such shareholder pressures without changing the legal norm defining the purpose of a business corporation would conflict with the fiduciary duties of corporate officers and directors, which are based on the traditional shareholder wealth maximization norm.87Marcel Kahan & Edward B. Rock, Systemic Stewardship with Tradeoffs, 48 J. Corp. L. 497, 500 (2023); see also Roberto Tallarita, The Limits of Portfolio Primacy, 76 Vand. L. Rev. 511, 564–65 (2023). In what follows we thus focus our analysis on a more ambitious version of PVM that includes changing the legal definition of corporate purpose to encompass the internalization of externalities that fall on other firms held in their shareholders’ portfolios.88Such an approach to PVM is precisely what motivated a 2022 class action lawsuit against Meta Platforms (formerly Facebook, Inc.), which alleged that the directors of Meta had breached their fiduciary duties by choosing to maximize the value of Meta rather the financial interests of Meta’s diversified shareholders. In particular, the complaint alleges that the directors failed to consider that shareholders with diversified portfolios may be subject to net losses in their portfolios due to Meta’s pursuit of a business model that maximizes its advertising revenue without regard to the harms this conduct inflicts on public health and, by extension, the value of diversified portfolios. See Complaint at 2, 18, 72, McRitchie v. Zuckerberg, No. 2022-0890 (Del. Ch. Oct. 3, 2022).

* * *

The main appeal of shareholder welfarism, in both its shareholder social preferences and portfolio value maximization guises, is that it seems to hold the promise of addressing the two key problems with conventional stakeholderism. First, it retains the basic norm that shareholder interests are primary in the management of a corporation. As such, shareholder welfarism might be compatible with the standard norms and incentives governing corporate affairs that put shareholders first, which the recent growth of institutional shareholders has further entrenched. Second, each form of shareholder welfarism provides a conceptual framework through which corporate management could determine, at least in principle, how to trade off among competing stakeholder interests. These two key aspects of the appeal of shareholder welfarism are shared by the ESV view. It too is compatible with existing norms that privilege shareholder interests and provides a clear objective to guide corporate management in trading off current profits in order to further stakeholder interests: long-term shareholder value.

IV.  EVALUATING THE THREE APPROACHES TO CSR THROUGH SHAREHOLDER GOVERNANCE

We now turn to evaluating the three approaches to pursuing corporate social responsibility through shareholder governance—enlightened shareholder value (“ESV”), shareholder social preferences (“SSP”), and portfolio value maximization (“PVM”)—based on their potential to induce the management of public companies to incur costs on a voluntary basis in ways that further the interests of other stakeholders in the firm (that is, to engage in CSR). We divide our analysis into three parts. We first evaluate the normative attractiveness of the corporate objective posited by each approach, ignoring the practical challenges to inducing corporate managers to pursue each objective. We focus simply on the extent to which each proposed corporate objective captures various social concerns about corporate behavior. We then turn to the feasibility of each approach in terms of the extent to which managers would have the information and incentives needed to pursue the posited corporate objective, taking as given the centralization of control in corporate managers. Finally, we consider the extent to which implementing shareholder welfarism by simply devolving more control to shareholders might improve corporate conduct.

A.  Normative Attractiveness of Each Corporate Objective

To what extent do the corporate objectives of ESV, SSP, and PVM capture CSR concerns? Our analysis in this Section can be thought of as adopting the assumption of no information costs and no agency costs: we imagine a world in which public companies fully maximize the corporate objective function posited under each approach. The corporate objective posited by the ESV view is long-term shareholder value, meaning the net present value of the future cash flows paid on the company’s equity, discounted based on the firm’s opportunity cost of capital.89Jensen, supra note 4, at 9. Note that long-term shareholder value is also a major component of the corporate objectives posited by SSP and PVM. ESV and its long-term shareholder value objective thus form a key benchmark against which to judge SSP and PVM. We begin by qualitatively characterizing the extent to which the long-term shareholder value objective of ESV fails to capture CSR concerns so that even in a world in which management was perfectly successful at maximizing long-term shareholder value in an enlightened way, there would remain significant residual social concerns. We then turn to the SSP and PVM objective functions and consider the extent to which the further considerations they incorporate in addition to long-term shareholder value might capture CSR concerns beyond the ESV baseline.

1.  Enlightened Shareholder Value

We begin by repeating an observation we made in our discussion of stakeholderism in Part I: corporate behavior is significantly shaped by the constraints and incentives produced by law and public policy, much of which is intended to address market failures and distributional concerns that arise from corporate conduct. This forms an important starting point for thinking about how, in a world in which managers perfectly maximize long-term shareholder value, there might remain social concerns about corporate conduct. Those concerns are, by definition, those not addressed by current law and public policy.

One category of social concerns about corporate conduct that would persist in such a world is with respect to the treatment of firm patrons. First, the outcomes for firm patrons—especially workers—might raise distributive concerns. Competitive labor markets, for example, operating under current tax and transfer policies induce a particular distribution of income and welfare in which low-skilled workers, in particular, earn income that many find unfairly low.90Thomas Piketty, Capital in the Twenty-First Century 304–35 (Arthur Goldhammer, trans., 2014). As we discussed above, maximizing long-term shareholder value generates some incentive for firms to pay their workers more than they otherwise would based on the value of incentivizing effort or firm-specific investment, but this is true only up to a point. Indeed, for workers for which such incentive contracting concerns do not loom large, the shareholder-value-maximizing wage might be little more than the competitive wage in the relevant labor market. Furthermore, it seems likely that such cases will often involve workers with relatively low levels of human capital whose low incomes raise the greatest distributive concerns from a social perspective. Put simply, efficiency wages and the like are no panacea for the standard concerns about the income inequality produced by market economies.

Another limitation of this class of ESV mechanisms stems from last period concerns. Firms have incentives to perform on implicit contracts in order to preserve the going concern value of the firm, which relies on the trustworthiness of the firm as perceived by current and future patrons. Implicit contracting thus depends critically on the firm and its patrons having a long future ahead of them. But as the probability that the firm will cease to operate and be liquidated goes up—due to business setbacks, for example—the incentives produced by the value of the firm’s reputation for trustworthiness are attenuated.

Market power of firms raises additional social concerns. While efficiency wage and implicit contracting considerations might moderate to some extent the incentive of shareholder-value-maximizing firms to exploit their market power, in the main, the long-term shareholder value objective is better understood as the key cause of the social problems posed by market power rather than as their solution.

In a similar way, maximizing long-term shareholder value provides no universal cure for other sources of contracting failures between the firm and various classes of firm patrons.91Indeed, the basic thesis of Henry Hansmann’s The Ownership of Enterprise is that such contracting failures can result in the efficient assignment of ownership of the firm being to a class of firm patrons other than investors. Hansmann, supra note 39, at 1–2. A firm that possesses better information than its customers about the safety of its products, for example, might well succumb to the temptation to cut back on safety to save costs, correctly concluding that the reputational and other costs of doing so are outweighed by the short-run savings even when viewed through the lens of long-term shareholder value.

With respect to externalities on non-firm patrons, the limits of ESV are even easier to see. By definition, when production or consumption of a firm’s output generates a negative technological externality, running the firm to maximize long-term shareholder value will result in socially excessive levels of the activity (and the reverse is true for positive externalities). The mechanisms discussed in Part II through which ESV can incentivize firms to improve their treatment of non-patrons do not change this powerful implication of economic theory. When externalities exist that are not effectively addressed through taxation or regulation, the private costs and benefits of the activity that drive the maximization of long-term shareholder value diverge from the social costs and benefits of the activity.

In summary, ESV mechanisms under the corporate objective of long-term shareholder value only mitigate and do not resolve social conflicts with respect to corporate conduct. We turn now to SSP and PVM to consider the extent to which the objective function posited by each might go further than ESV in motivating CSR.

2.  Shareholder Social Preferences

The corporate objective under the SSP view is based on two key components of shareholders’ well-being: (1) the long-term value of the shares and (2) shareholders’ social preferences with respect to corporate conduct. The weight each shareholder puts on these two components depends on their own preferences. As well, the specific content of shareholders’ social preferences will vary from shareholder to shareholder. To calculate aggregate shareholder welfare, individual shareholders’ well-being levels are weighted by their share ownership and summed.

Because of heterogeneity across shareholders in the strength and content of their social preferences, aggregate shareholder welfare for a corporation will depend on who owns the shares of the company. In turn, the decisions of individuals to hold the shares may well depend on the conduct of the corporation and the social preferences of the individuals. For now, we adopt the simplifying assumption that all shareholders are fully diversified, so that there is no variation in the share-weighted social preferences of shareholders of different public companies.92We consider the sorting of shareholders across firms infra Section IV.B.2.i.

Under these assumptions, how would maximizing shareholder welfare, taking into account the social preferences of shareholders, change corporate conduct relative to maximizing long-term shareholder value? Consider first the weight that aggregate shareholder welfare would put on long-term shareholder value. This is an empirical question based on the share-weighted preferences of corporate shareholders. But we make three points that together point to the conclusion that aggregate shareholder welfare would be largely, perhaps even overwhelmingly, based on long-term shareholder value rather than shareholders’ social preferences.

To begin, it is useful to contrast shareholder welfare with overall social welfare. Social welfare does include as a component a firm’s long-term shareholder value—the well-being of the claimants to that value count, of course, in any appropriate measure of social welfare. But social welfare also includes the well-being of those who are not shareholders of the firm. In contrast, shareholder welfare would put weight on non-shareholders’ well-being based only on shareholders’ social preferences. Unless shareholders were perfectly altruistic in the sense that their preferences put as much weight on others as on themselves, this results in shareholder welfare putting greater relative weight on firm value than does social welfare. This effect alone means that maximizing shareholder welfare will generally not provide an incentive for managers to sacrifice profits to the extent required for the firm to behave appropriately as a social matter. Consider, for example, a profitable factory that emits such a large amount of pollution that, from a social welfare perspective, it should be shut down. Because shareholder welfare puts much more weight on firm profits than social welfare does, it will often not be in shareholders’ interests in such a situation to shut down the plant even including consideration of their social preferences.

Second, the shareholders of a public corporation are insulated from the social and moral pressures that generate other-regarding behavior at the individual level.93Elhauge, supra note 41, at 758–59. This is due in part to the complex governance structures that stand between individual shareholders and corporate decision-making that make shareholders anonymous to those who might impose social sanctions for harm done by the corporation as well as due to diversified shareholders’ basic lack of information about corporate affairs (ignorance is bliss).94Id. at 798. Einer Elhauge argues that this insulation will result in shareholders putting much more weight on corporate profits relative to social concerns than would sole proprietors, who are far less insulated.95Id. at 799. This is even more strongly the case with respect to shareholders who own interests in corporate shares through intermediaries like mutual funds and are therefore “double insulat[ed].”96Id. at 817. In sum, from a revealed preference perspective, the welfare of diversified shareholders might be understood as stemming overwhelmingly from shareholder value rather than from social preferences.

Finally, what little weight shareholder welfare does put on social concerns as opposed to shareholder value is further muted by conflicts among shareholders about social issues. Hart and Zingales introduce the idea of shareholder welfare in a simple model in which the social concern is about pollution that is a by-product of firm operations and shareholders’ preferences vary only in terms of the weight they put on environmental harm from the firm’s pollution versus on their own wealth.97Hart & Zingales, supra note 71, at 252–53. In this framework, aggregate shareholder welfare will be based on the share-weighted average of the weights individuals put on environmental harm relative to personal wealth.

But corporate activities typically pose trade-offs not just between profits and social concerns but also among competing social concerns. As a result, conflicts among shareholders in their views on social issues effectively further reduce the weight of shareholder preferences in determining what maximizes shareholder welfare. In some cases, these conflicts are direct. Consider abortion or affirmative action. Some socially minded investors want less of these things; some want more. In those cases, the competing social preferences of different shareholders cancel out to some extent so that, on net, shareholder social preferences get less weight in determining shareholder welfare.

But even for social issues that nobody is against per se, like clean air or good jobs, there are often indirect conflicts stemming from shareholders’ social preferences. Consider a manufacturing firm that causes pollution as a by-product of its production process but also provides jobs in a community with scarce economic opportunities.98See Alperen A. Gözlügöl, The Clash of ‘E’ and ‘S’ of ESG: Just Transition on the Path to Net Zero and the Implications for Sustainable Corporate Governance and Finance, 15 J. World Energy L. & Bus. 1, 4 (2022) (arguing that the transition to net-zero greenhouse gas emissions will result in certain regions suffering substantial employment losses). The choice of scale of the firm’s output poses trade-offs between environmental quality and jobs. As a result, a socially minded investor who cares about both might ultimately prefer a level of output little different from the profit-maximizing level of output. In contrast, shareholders who care more about the environment than jobs might prefer a lower level of output, and vice-versa for a shareholder more concerned about jobs. The median shareholder’s preferences might then be close to the profit-maximizing level of output. So these indirect conflicts about social issues also, in effect, further mute the role of social preferences in shareholder welfare and increase the role of long-term shareholder value.

Note as well that corporations do not generally face binary decisions—like either protect the environment or preserve jobs—but rather face a continuum of choices, as in the example of a firm’s choice of level of output. As a result, having a bare majority of shares held by shareholders who lean in one direction on such trade-offs—toward the environment, say—does not mean that the conflicting preferences of the remaining shareholders do not matter for determining the operational decision that maximizes shareholder welfare. For a firm facing a continuum, or at least a large number, of potential operational decisions, the presence of a significant minority of shareholders who care more about jobs than the environment will pull the shareholder-welfare-maximizing choice in the direction of preserving jobs and away from protecting the environment.99We put to the side here more profound complications posed by conflicts among preferences of individuals for aggregating those preferences to a social choice, for example, the possibility that majority voting over choices might fail to yield a stable outcome. See generally Kenneth J. Arrow, Social Choice and Individual Values (1951).

In light of these considerations, the social issues for which incorporating shareholders’ social preferences into the corporate objective might potentially make a meaningful difference, relative to the ESV baseline, in motivating CSR would generally be issues on which there is a broad and strong social consensus. But these are exactly the set of issues for which the residual social concerns left under the ESV approach after fully maximizing long-term shareholder value are likely to be minimal, for two reasons.

First, social issues for which there is a strong social consensus are much more likely to be effectively addressed by law and public policy. Federal and state law, for example, provide powerful controls on corporate conduct to address many social concerns raised by corporate operations, from the safety of motor vehicles, to the health consequences of tobacco consumption, to the emission of particulate matter by industrial activities. Our claim is most certainly not that the political process is perfect or that current public policy fully addresses all social concerns about corporate conduct. Rather, it is that the specific issues for which there is sufficient social consensus such that the social preferences of shareholders form a meaningful component of shareholder welfare are precisely the issues that are most likely to be effectively addressed by public policy. Indeed, corporate shareholders’ preferences put less weight on average on the social concerns raised by corporate conduct than does the overall polity, for reasons given above.  It thus seems likely that for many issues for which there is a strong social consensus, public policy will go well beyond what the company’s shareholders would prefer in reining in corporate conduct.100An example of this dynamic can be seen in the Rule 14a-8 campaign by environmentally oriented shareholders such as As You Sow against oil production companies between 2017 and 2019. These shhareholders sought to compel greater corporate disclosure regarding methane gas leaks arising from their oil production operations. See, e.g., Dominion Energy, Inc.: Request for Report on Methane Leaks, As You Sow (Jan. 31, 2018), https://www.asyousow.org/resolutions/2018/01/31/dominion-energy-inc-request-for-report-on-methane-leaks [https://perma.cc/XTR2-BU2Q]. Public polls at this time suggested that 74% of respondents “strongly support[ed]” or “somewhat support[ed]” regulations to reduce methane gas leaks, see Climate Nexus, National Poll Toplines (2021), https://climatenexus.org/wp-content/uploads/2015/09/Climate-Nexus-National-Poll-2021-Methane-Infrastructure-Toplines.pdf [https
://perma.cc/5X4H-UE2D], which may explain why the Biden-Harris administration implemented its Methane Emissions Reduction Action Plan in 2022, see Fact Sheet: Biden Administration Tackles Super-Polluting Methane Emissions (Jan. 31, 2022), https://www.whitehouse.gov/briefing-room/statements-releases/2022/01/31/fact-sheet-biden-administration-tackles-super-polluting-methane-emissions [https://
perma.cc/QFT9-8GCN]. Notably, despite the widespread public support for regulating methane leaks, shareholder support for 14a-8 proposals aimed at enhancing methane leak disclosures, while occasionally reaching 50% support, often drew far less than majority support. See Shareholders Are Plugging Methane Leaks Themselves, As You Sow (June 1, 2018), https://www.asyousow.org/blog/2018/6/1/shareholders-are-plugging-methane-leaks-themselves [https://perma.cc/9W2A-26N5].

Second, the broad social consensus we are supposing would include not just shareholders but also other classes of firm patrons, including its workers, managers, and customers. The social preferences of firm patrons can provide strong shareholder value reasons for the firm to act in ways that are consistent with those social preferences. Failing to do so risks inviting a backlash from these other classes of firm patrons that might have major financial consequences.101Barzuza et al., supra note 65, at 265. As BlackRock’s CEO Larry Fink put it in his 2022 letter to CEOs, “Employees need to understand and connect with your purpose; and when they do, they can be your staunchest advocates. Customers want to see and hear what you stand for as they increasingly look to do business with companies that share their values.” Larry Fink, Larry Fink’s 2022 Letter to CEOs: The Power of Capitalism, BlackRock (2022), https://www.blackrock.com/corporate/investor-relations/
larry-fink-ceo-letter [https://perma.cc/C82G-E8DM].

Consider, for example, explicit and open racism in a firm’s treatment of its customers. A recent episode involving Starbucks is instructive. In 2018, a Starbucks employee called the police after two Black men entered a Starbucks in Philadelphia and sat down without purchasing anything and, when store employees asked them to leave, declined to do so. The police forcibly removed the men, leading to national headlines, a public apology by the Starbucks CEO, and the hashtag #BoycottStarbucks trending on Twitter.102Matt Stevens, Starbucks C.E.O. Apologizes After Arrests of 2 Black Men, N.Y. Times (Apr. 15, 2018), https://www.nytimes.com/2018/04/15/us/starbucks-philadelphia-black-men-arrest.html [https
://perma.cc/FGK9-Z5AA].
No reference to Starbucks shareholders’ social preferences is needed to explain the decision by Starbucks management several days later to close 8,000 stores to conduct racial bias training of employees.103Rachel Abrams, Starbucks To Close 8,000 U.S. Stores for Racial-Bias Training After Arrests, N.Y. Times (Apr. 17, 2018).

In summary, under the SSP shareholder welfare objective, it is long-term shareholder value that is the key driver of decisions to incur costs to further stakeholder interests, not the social preferences of shareholders, which are conflicted, muted, and often prefer less protection of stakeholder interests than provided by law.104In contrast, Broccardo et al. argue that diversified shareholders, in casting votes about corporate issues, will put more weight on social concerns than a sole proprietor would since each shareholder bears only a fraction of the costs of the firm behaving more responsibility. See Broccardo et al., supra note 78, at 3103. We discuss Broccardo et al.’s model in more detail infra Section IV.C.

3.  Portfolio Value Maximization

The corporate objective under the PVM approach is diversified shareholders’ portfolio value. To evaluate its normative desirability, we maintain for now the simplifying assumption that all investors are fully diversified—that is, they hold the market portfolio of all investible risky assets with each asset weighted in proportion to its value. This is in fact a key assumption underlying the standard model of valuation managers are taught in MBA programs, which is based on the Capital Asset Pricing Model (“CAPM”).105See, e.g., Richard A. Brealey, Stewart C. Myers & Franklin Allen, Principles of Corporate Finance 185–99 (10th ed. 2011). CAPM provides the original intellectual foundations for the specific model of financial management by which managers are supposed to pursue long-term shareholder value. We begin by sketching how that model works in order to frame more precisely how the PVM approach proposes managers should deviate from it.

In the standard model of corporate decision-making, diversified shareholders want managers to follow the “NPV Rule”: invest in every project that has a positive net present value (“NPV”).106See id. at 101–03. The NPV of a project is calculated by converting (“discounting”) all of the future cash flows associated with the project to their present value and then summing those present values as follows:

,        (1)

where  is the net cash flow received from the project in period T and r is the risk-adjusted discount rate for the project.

The assumption of CAPM—that all investors are optimally diversified—plays a key role in the determination of the appropriate discount rate.107For a textbook treatment of CAPM, see id. at 185–203. To capture the cost to investors of bearing the risk of the project, a “risk premium” is added to the risk-free rate (typically taken to be the return on government obligations) to arrive at the risk-adjusted discount rate. But crucially, CAPM considers the risk of a project from a portfolio perspective. That is, a project’s risk is measured not in terms of the degree of uncertainty of the project’s cash flows considered in isolation but rather in terms of the increment in portfolio risk if the project were added to a diversified portfolio. This matters because one component of a project’s risks—the idiosyncratic component—disappears when the project is held in a diversified portfolio. A diversified investor only has to be compensated for bearing the risks that they actually have to bear, which is the undiversifiable, systematic component of a project’s risk. In CAPM, the only source of systematic risk comes from the correlation between a project’s cash flows and the overall market return, which is referred to as the project’s beta. The standard shareholder value approach thus already adjusts the denominators of the fractions in the above expression for NPV based on a portfolio perspective. So the idea that corporate managers should take a portfolio perspective on the interests of shareholders is actually an old one and entirely conventional. It forms a core component of standard shareholder value theory.

The PVM approach, however, pushes this portfolio perspective further. It incorporates into the cash flows of the project not just the cash flows received by the firm but also the increment in cash flows paid on any other securities in the market portfolio. This entails adjusting not only the denominators of the terms in the expression for NPV, but also their numerators. The resulting NPV expression under the PVM approach is:

        (2)

The numerators in the PVM-modified expression for NPV include both the expected cash flows from the project that will accrue to the instant corporation (the ’s) as well as the spillover expected cash flows for other securities resulting from externalities (the ’s), which could be on net either positive or negative in any given period. For most corporate decisions, the bulk of the cash flows at the market portfolio level in fact accrue to the securities issued by the corporation making the decision. The question we grapple with in this Section is the extent to which the consideration of the additional cash flows to other portfolio securities that the PVM approach requires—assuming no agency costs or information problems—will motivate CSR beyond that justified on the basis of maximizing long-term shareholder value under ESV. We reach an even more negative conclusion than the one we reached in evaluating the SSP objective function: the portfolio value objective will not only produce little additional motivation for CSR, but it will also provide new motivations for socially destructive corporate conduct.

First, taking a portfolio perspective on expected cash flows produced by corporate decisions captures only a small portion of the technological externalities of corporate conduct since the bulk of such externalities fall on interests that are not part of the market portfolio. These interests include the health and well-being of consumers as well as the interests of producers that are not owned in the market portfolio.108The aggregate portfolio of the stockholders of a public company would include some assets that are not public securities, and in principle the PVM objective function could include the value of those additional assets. However, we assume that for most investors in public companies, their portfolios are dominated by securities issued by publicly listed firms and other publicly available investments, like U.S. Treasury securities. For instance, even for an extremely diversified institutional investor such as CalPERS, well over half of its $462 billion of assets under management consists of global public equity and publicly offered investment securities such as investment grade debt and U.S. Treasury securities. See CalPERS, Trust Level Review 13, 34 (2023), https://www.calpers.ca.gov/docs/board-agendas/
202309/invest/item05b-01_a.pdf [https://perma.cc/98NR-PMPF]. As a result, the PVM objective function would largely fail to capture external effects on other kinds of assets.

To be concrete, consider the facts alleged in Aguinda v. Texaco, a class action filed on behalf of residents of certain regions of Ecuador and Peru to recover for property damage, personal injuries, and increased risk of disease allegedly caused by Texaco’s improper waste disposal practices in its oil extraction operations in Ecuador.109Aguinda v. Texaco, Inc., 142 F. Supp. 2d 534, 537 (S.D.N.Y. 2001). The plaintiffs alleged that Texaco engaged in a range of wrongful conduct, including dumping large quantities of toxic by-products of the drilling process into local rivers and landfills.110Jota v. Texaco, Inc., 157 F.3d 153, 155 (2d Cir. 1998) (consolidated on appeal with Aguinda v. Texaco, 945 F. Supp. 625 (S.D.N.Y. 1996)). Texaco allegedly did this to save money, netting additional profits of $500 thousand to $1 million per well.111Class Action Complaint at 19, Ashanga Jota et al. v. Texaco, Inc, No. 94 Civ. 9266 (S.D.N.Y. Dec. 28, 1994). The pollution released by Texaco poisoned the local ecosystem, causing environmental harm, economic losses to local fishermen and agriculture, and serious injuries and disease among local residents.112Id. at 5–13.

These allegations represent a paradigmatic case of socially harmful corporate behavior that CSR advocates hope to address. The harms suffered by local residents constituted negative technological externalities that were not effectively controlled through regulation or private law remedies.113The class actions brought seeking damages and equitable relief in U.S. courts were ultimately dismissed on the basis of forum non conveniens. Aguinda v. Texaco, Inc., 303 F.3d 470, 473–74, 480 (2d Cir. 2002). But they also illustrate a key limitation of the PVM approach: hardly any of these externalities would have manifested as reductions in expected cash flows to securities in the market portfolio. To be sure, the kinds of costs at issue in this example—costs to human health, ecosystems, and small-scale producers—might ultimately have second-order effects on companies in the market portfolio as, for example, the resulting shifts in supply and demand in various markets affect prices of companies’ inputs and outputs. But those effects on companies are de minimis and, for that matter, could be on net positive if, for example, the resulting fall in production by small-scale producers resulted in a reduction in supply of products sold by larger companies. To a first approximation, the ’s for this project would be zero, despite the sizable social externalities at issue.114A similar evidentiary challenge appears with regard to the McRitchie v. Zuckerberg class action. See Complaint, supra note 88, at 74. The technological externality at the heart of the case relates to the alleged costs of Meta’s pursuit of advertising revenue on public health and the rule of law and, by extension, economic growth. Even assuming Meta’s operations created these externalities, it is far from clear whether its actions would have adversely affected a diversified investor’s portfolio value, absent an express netting of the costs and benefits of Meta’s efforts to maximize the value of the company.

This is also true for larger-scale externalities. Consider climate change, which has been aptly described as “the mother of all externalities.”115Richard S. J. Tol, The Economic Effects of Climate Change, 23 J. Econ. Persps. 29, 29 (2009). Essentially every business project produces some amount of greenhouse gas emissions, the accumulation of which in the atmosphere leads to warming of the planet over time. Climate change is expected to cause a manifold set of impacts on human well-being. The most recent report by the Intergovernmental Panel on Climate Change (“IPCC”) provides a useful taxonomy of the ways climate change is expected to affect human systems:

  1. Impacts on water scarcity and food production.
  2. Water scarcity.
  3. Agriculture / crop production.
  4. Animal and livestock health and productivity.
  5. Fisheries yields and aquaculture production.
  6. Impacts on health and wellbeing.
    1. Infectious diseases.
    2. Heat, malnutrition and other.
    3. Mental health.
    4.  
  7. Impacts on cities, settlements and infrastructure.
    1. Inland flooding and associated damages.
    2. Flood / storm induced damages in coastal areas.
    3. Damages to infrastructure.
    4. Damages to key economic sectors.116IPCC, 2022, Climate Change 2022: Impacts, Adaptation and Vulnerability: Contribution of Working Group II to the Sixth Assessment Report of the Intergovernmental Panel on Climate Change 10 (H.-O. Pörtner, D.C. Roberts, M. Tignor, E.S. Poloczanska, K. Mintenbeck, A. Alegría, M. Craig, S. Langsdorf, S. Löschke, V. Möller, A. Okem & B. Rama, eds., 2022), https://report.ipcc.ch/ar6/wg2/IPCC_AR6_WGII_FullReport.pdf [https://perma.cc/
      8T33-BXXS].

While some of these categories, especially those under “[i]mpacts on cities, settlements and infrastructure,” would include substantial effects on companies in the market portfolio, this taxonomy reveals that the scope of the harms from climate change is far broader than those effects.

Indeed, the United Nations Environment Programme’s Finance Initiative (“UNEP FI”) developed a methodology for assessing the impact of climate change on the portfolios of institutional investors that illustrates the relatively small portion of the costs of climate change that affect the value of the market portfolio in May 2019.117United Nations Env’t Programme Fin. Initiative, Changing Course: A Comprehensive Investor Guide to Scenario-Based Methods for Climate Risk Assessment, in Response to the TCFD (2019), https://www.unepfi.org/wordpress/wp-content/uploads/2019/05/TCFD-Changing-Course-Oct-19.pdf [https://perma.cc/R5BD-7HFT]. The physical risks from climate change included in the analysis are limited to asset damage and business interruption from extreme weather events, a fairly small component of the myriad social costs of climate change identified by the IPCC.118Id. at 16. Physical risks are what economists would consider the social costs of climate change, including all effects on human society described in the IPCC 2022 report summarized above. Transition risks, on the other hand, refer to business issues raised by the shift from a high-carbon economy to a low-carbon economy induced by technological change and government policy. For example, the risk that an oil company’s proven reserves would fall in value due to the imposition of a carbon tax or fall in demand for oil would constitute a transition risk but should not be considered a social cost of climate change in an economic sense. The PVM approach aspires to induce companies to internalize the physical risks posed by greenhouse gas emissions. Tallarita, supra note 87, at 517. This reflects how limited a perspective the PVM objective function brings to the social costs of even large-scale externalities like climate change.

A similar issue concerns the geographic distribution of the social costs of climate change. Existing estimates show that the costs of climate change will be disproportionately borne by lower income regions. For instance, Africa and India are estimated to have aggregate climate damages as a percentage of GDP that are nearly 800% and 1,000%, respectively, greater than those estimated for the United States.119William D. Nordhaus & Joseph Boyer, Warming the World: Economic Models of Global Warming 91 (2000). Yet to the extent investors access the market portfolio by means of investing in public securities and private company debt, the market portfolio of these securities is tilted toward economic activity in North America and Europe. The standard measure of the extent to which a country’s economic activity occurs through public companies and private debt is the country’s market-cap-to-GDP ratio. In general, the GDP ratio is much higher for developed economies like those in North America and Europe that are relatively less exposed to the costs of climate change than the GDP ratio for the developing economies that face the largest risks.120See Martin Čihák, Asli Demirgüč-Kunt, Erik Feyen & Ross Levine, Financial Development in 205 Economies, 1960 to 2010, J. Fin. Persps., July 2013, at 1, 7. The authors include the value of both public equity and debt and private debt in the numerator of this ratio. Id.

This geographic mismatch problem also raises difficulties for one of the standard methodologies for estimating the degree to which reductions in carbon emissions would increase diversified investors’ portfolio values. For instance, in an influential paper in Nature Climate Change, Simon Dietz, Alex Bowen, Charlie Dixon, and Philip Gradwell estimate that, relative to a world without climate risk, investors can expect to lose $2.5 trillion due to the impact of climate risk on global financial assets.121Simon Dietz, Alex Bowen, Charlie Dixon & Philip Gradwell, ‘Climate Value at Risk’ of Global Financial Assets, 6 Nature Climate Change 676, 678 (2016). Madison Condon likewise estimates that if BlackRock could induce Chevron and Exxon to cut industrial emissions such that 1% of industrial emissions were removed each year through 2100, the global reduction in climate damages would have a net present value of $385 billion.122Condon, supra note 85, at 46 n.237. Given the size of BlackRock’s portfolio, she estimates that BlackRock would therefore avoid damages to its portfolio with a net present value of $9.7 billion, which would be sufficient to offset BlackRock’s losses in the equity values of Chevron and Exxon.123Id. But to arrive at these estimates, these scholars all utilize William Nordhaus’s Dynamic Integrated Climate-Economy (“DICE”) model to estimate the impact of climate change on global GDP growth.124See Dietz et al., supra note 121, at 677; Condon, supra note 85, at 46. They then assume that climate change will have a proportional effect on global financial assets given past research showing that aggregate financial returns generally track GDP growth.125See Dietz et al., supra note 121, at 676; Condon, supra note 85, at 46 n.237. A further problem with Condon’s analysis is that she uses the wrong denominator for the fraction of climate change impacts internalized by BlackRock’s portfolio under management. Formally, Condon first estimates the present value of the reduction in climate damages on global GDP and then assumes that the value of the damage reduction to BlackRock is based on BlackRock’s share of the global economy based on the ratio of BlackRock’s assets under management ($7.43 trillion) to global GDP (roughly $80 trillion). Condon, supra note 85, at 2 n.3, 46 n.237. Because global GDP is a measure of income, the relevant denominator for this purpose should be global financial assets or roughly $143.3 trillion according to Dietz et al. Dietz et al., supra note 121, at 678. Using the correct denominator, the estimated reductions in damages to Blackrock’s portfolio would decline from $9.7 billion to about $5.5 billion, which is less than the $6.5 billion that Condon estimates BlackRock would lose due to declines in the equity values of Chevron and Exxon. Condon, supra note 85, at 46. However, the DICE model integrates the heterogeneous effects of climate change on different countries to produce a single estimate of the effect of climate change on global GDP growth, ignoring the fact that the costs of climate change will not be shared equally across all countries. This methodology therefore overestimates the effect of climate change on the growth rate for the market portfolio, which is tilted toward economic activity in North America and Europe.

As noted by Roberto Tallarita, a related issue with the objective function of PVM is that it discounts future costs and benefits using the opportunity cost of capital.126Tallarita, supra note 87, at 548–54. But for costs and benefits that play out over long time scales that span generations, like those of climate change, economists typically apply a discount rate that is much lower than the opportunity cost of capital to account for intergenerational distributional considerations.127Moritz A. Drupp, Mark C. Freeman, Ben Groom & Frikk Nesje, Discounting Disentangled, 10 Am. Econ. J.: Econ. Pol’y 109, 112–13 (2018). This results in the PVM approach massively undercounting the costs of climate change, most of which will not accrue for many decades.128Tallarita, supra note 87, at 548–54.

To give a rough numerical sense for the magnitude of this issue, note first that the present value of the future costs of climate change, when using social discount rates in the range typically used for climate policy, stems largely from impacts that will occur beyond the year 2200.129Nicholas Stern, The Economics of Climate Change, 98 Am. Econ. Rev. 1, 20 (2008). For example, in Stern’s influential The Economics of Climate Change, 90% of the present value of the social costs of carbon emissions stem from impacts that occur after 2200. Id. To simplify, suppose that all of those impacts occurred in 2200, which is 177 years from the year 2023. Suppose that the right social discount rate to use to convert those costs to present value is 2%, a number often used by experts.130Drupp et al., supra note 127, at 128 (reporting that the median social discount rate recommended by experts is 2%). In a 2022 analysis of the social costs of carbon, the EPA similarly used 2% as its central discount rate target. See Env’t Prot. Agency, Supplemental Material for the Regulatory Impact Analysis for the Supplemental Proposed Rulemaking, “Standards of Performance for New, Reconstructed, and Modified Sources and Emissions Guidelines for Existing Sources: Oil and Natural Gas Sector Climate Review” 2 (2022), https://www.epa.
gov/system/files/documents/2022-11/epa_scghg_report_draft_0.pdf [https://perma.cc/P7C9-SW55].
At that social discount rate, each dollar of future climate change costs should be discounted by the factor 1/1.02177, which comes out to 0.03. A $1 trillion future climate change cost in 2200 would then be considered worth $30 billion in present value terms. But applying the 12% real discount rate typically used by corporate managers, the PVM approach would use a discount factor of just 1/1.12177 or 0.000000002. Under the PVM approach, that $1 trillion future social cost of climate change comes out to just $1,943 in present value terms. Or in different terms, the PVM approach would capture only the fraction (1/1.12177)/(1/1.02177) or 0.00000007 of the present value of the costs of climate change in 2200 (and even less of those beyond). Even if managers used a much lower discount rate of 7% under PVM, this fraction still comes out to just 0.0002. Discounting alone thus results in the PVM objective function internalizing only a trivial fraction of the social costs of climate change.

The UNEP FI report also illustrates another conceptual problem with the PVM approach: the methodology incorporates the positive business opportunities created by climate change for companies in the market portfolio.131United Nations Env’t Programme Fin. Initiative, supra note 117, at 44–45. The transition to a low-carbon economy and adaptation to a warming planet will require investment in technologies and infrastructure in a range of sectors. To give one example, consider a concrete seawall installed in New York Harbor to address storm surges caused by climate change. The Army Corps of Engineers has proposed the construction of such a barrier at a cost of some $119 billion.132Anne Barnard, The $119 Billion Sea Wall that Could Defend New York . . . or Not, N.Y. Times (Aug. 21, 2021), https://www.nytimes.com/2020/01/17/nyregion/the-119-billion-sea-wall-that-could-defend-new-york-or-not.html [https://perma.cc/PTK7-SE4A]. If such a seawall were built in order to deal with climate change, it would count as among the negative externalities of climate change—it is a real resource use caused by the warming of the planet. But from a PVM perspective, the construction of a seawall represents an enormous business opportunity. In other words, while the aspiration of the PVM approach is to incorporate such costs as negative adjustments to expected cash flows for business projects that contribute to climate change (that is, negative ’s in the PVM-adjusted NPV expression above), in fact faithful application of the PVM approach would incorporate them at least in part as positive adjustments since the construction of the seawall will produce profits for companies in the market portfolio (that is, as positive ’s).

A final problem with the PVM objective function’s treatment of technological externalities is with respect to its interaction with public policies designed to address such externalities. Consider, for example, a pollution externality caused as a by-product of a certain production process, and suppose the externality is addressed at the public policy level with a Pigouvian tax set at the marginal social cost of the externality. As a result, the private profit-maximization problem facing firms that emit that form of pollution mirrors the social problem of choosing efficient behavior. But consider what would happen if managers of the polluting firms were instead to set firm policy following the PVM approach. Those managers would consider not only the Pigouvian tax but also the portion of the externality that reduced the value of other firms in the portfolio so that a portion of the externality would be double counted. As a result, they would, at the margin, be over-deterred from producing pollution. In short, the PVM approach, unlike ESV, does not integrate well with public policy approaches to addressing externalities.133This problem could be mitigated, in principle, by calibrating the level of the Pigouvian tax to be equal to the portion of the externality that falls on interests other than securities in the market portfolio. However, it is not clear how policymakers could determine that amount.

In contrast to these failures with respect to technological externalities, the PVM approach is far better suited to capture pecuniary externalities. One reason is that pecuniary externalities largely involve a company’s competitors, a significant fraction of which are public companies. Consider the airline industry, which is dominated by public companies.134See Niraj Chokshi, Frontier Airlines I.P.O. Signals a Travel Industry Recovery, N.Y. Times (June 15, 2021), https://www.nytimes.com/2021/04/01/business/frontier-airlines-ipo.html [https://perma
.cc/W2N3-BCYT] (noting that as of 2021, the ten largest airlines in the U.S. are publicly listed).
When Delta Airlines cuts its fares on the D.C.-Boston route and gains market share, it reduces the value of its competitors on that route, which are largely public companies. As we noted above, however, this feature of PVM is really a bug. If companies fully maximized diversified investors’ portfolio value, the resulting reduction in competition would harm consumers and workers even as it benefited investors. The PVM objective function thus poses significant harms to firm patrons relative to the ESV baseline.

To summarize, the objective function under PVM is socially perverse. It fails to capture effectively much of the technological externalities produced by corporate activities while at the same time having the potential to produce a form of market power that would be socially destructive to firm patrons. By our lights the PVM objective function is unattractive as a normative matter.

B.  Feasibility for Corporate Managers

We now consider whether managers would have the information and incentives they would need to pursue the stated corporate objective under each approach. We begin by reiterating the insight that both SSP and PVM effectively build on ESV since long-term shareholder value is a primary component of both shareholder welfare and portfolio value. As such, we first evaluate the information and incentive problems that might confound implementing long-term shareholder value as the corporate objective under ESV. Having established these problems as a baseline, we then turn to analyzing SSP and PVM. In this section we take as fixed the centralization of management of the corporate form in the board of directors and hired professional managers. We analyze the extent to which changing the legal and business norm on the objective of a business corporation from the long-term shareholder value objective of ESV to either the SSP or PVM objectives would improve corporate behavior given the resulting incentives and information of corporate managers. We then consider in Section IV.C whether a structural change to corporate control that would give shareholders a greater say in operational decision-making, as some advocates of shareholder welfarism have urged, would be likely to improve corporate behavior.

1.  Enlightened Shareholder Value

i.  Information

The informational burden of ESV is considerable. Part of the challenge stems from the inevitable uncertainty with respect to contingencies far out in the future. As we have emphasized, ESV arguments for CSR often have a temporal structure in which the company incurs costs in the near term in order to achieve benefits to stockholders that play out over a long period into the future. Consider, for example, investing in renewable energy, shutting down a dirty factory, or auditing the supply chain for safe labor practices. To what extent would sacrificing corporate profits in those ways today enhance shareholder value over the long-term?

While these questions are no doubt complicated, we view the information gathering and analytic challenges posed by ESV as squarely in the wheelhouse of corporate management. First, the intertemporal structure typical of ESV is not unique but rather is standard fare in business management. Corporate managers face similar intertemporal challenges in many other aspects of business strategy unrelated to CSR. Should the firm expand production? Should it invest more in research and development? Does it have the optimal capital structure? Business schools train managers in analytic techniques—most prominently discounted cash flow analysis—to grapple with such ubiquitous trade-offs and uncertainties entailed by managing a business.

Today, the specific strategic issues raised by CSR under the ESV approach are part of the bread-and-butter of business school curriculums. New York University’s Stern School of Business, for example, currently offers no fewer than 33 courses under the “Sustainable Business and Innovation” specialization, including course titles such as “Corporate Branding & Corporate Social Responsibility,” “Sustainability for Competitive Advantage,” and “Sustainable Capitalism: A Longer Term Finance Perspective.”135Course Index, NYU Stern Sch. of Bus., https://www.stern.nyu.edu/programs-admissions/
full-time-mba/academics/course-index [https://perma.cc/725J-N7FH]. By comparison, a mere thirteen courses are offered at NYU under the “Real Estate” specialization. Id. Not to be outdone, UC Berkeley’s Haas School of Business maintains the Institute for Business and Social Impact which oversees three separate centers focused on corporate sustainability and curates the Michaels Graduate Certificate in Sustainable Business. Institute for Business & Social Impact, Berkeley Hass, https://haas.
berkeley.edu/responsible-business/curriculum [https://perma.cc/3ZSV-DM5B]. MBA students at Haas can choose from twenty-nine courses focused on corporate sustainability such as “Climate Change and Business Strategy,” “Business and Sustainable Supply Chains,” and “Strategic and Sustainable Business Solutions.” Id.
From the course catalogs alone, it is clear that ESV is a major part of the analytic tool kit and worldview imparted to MBA students. Indeed, business school professors are among the most vociferous proponents of ESV.136See, e.g., Edmans, supra note 43, at 55–56.

Stock prices provide an additional source of information for a manager trying to understand the long-term value generated by current corporate policies. Stock markets incentivize the production and aggregation of information about corporate value by stock traders. Even if a manager is concerned that stock prices do not fully reflect long-term value, stock prices surely provide some relevant information to management regarding how to maximize long-term value. For example, the fact that Tesla and General Motors trade today with price-to-earnings ratios of 70 and 4, respectively, must say something about the future of internal combustion engines.137Tesla Inc., Google Fin., https://www.google.com/finance/quote/TSLA:NASDAQ [https://
perma.cc/QLF5-T9J2]; General Motors Co., Google Fin., https://www.google.com/finance/quote/
GM:NYS [https://perma.cc/U7UP-KS5Z].

In summary, while maximizing long-term shareholder value under ESV puts a substantial informational burden on corporate management, there are good reasons to believe that managers are able to assemble and process a great deal of information about how best to further stakeholder interests so as to maximize long-term shareholder value.

ii.  Incentives

Although ESV strikes us as substantially feasible from an information perspective, the story is more complicated with respect to managers’ incentives. As discussed in Part I, one reason for optimism stems from the structure of corporate law, which is generally designed with the goal of incentivizing management to maximize long-term shareholder value. Furthermore, Delaware courts have required corporate boards to put in place information and reporting systems designed to safeguard against risks to the company’s stakeholders that might ultimately harm shareholder interests through, for example, sullying the company’s reputation.138See, e.g., Marchand v. Barnhill, 212 A.3d 805, 809 (Del. 2019)(declining to dismiss a complaint charging a company’s board with breaching its fiduciary duties by failing to implement a monitoring system for food safety and observing that the company could only thrive if its customers “were confident that its products were safe to eat”).

Executive compensation for senior officers also produces substantial incentives for managers to maximize shareholder value. Much of these incentives stem from the significant equity component of managers’ pay packages, which directly links the wealth of managers to the wealth of shareholders. For example, for the median CEO of an S&P 500 firm as of 2011, a 1% increase in the value of the company’s shares would produce an increase in the wealth of the CEO of about $500,000 due to their holdings of company stock and stock options.139Kevin J. Murphy, Executive Compensation: Where We Are, and How We Got There, in 2A Handbook of the Economics of Finance 211, 236–37 (George M. Constantinides, Milton Harris & Rene M. Stulz eds., 2013).

Yet, while corporate governance is very much oriented toward the long-term shareholder value corporate objective of ESV, by no means does our corporate system produce perfect incentives for corporate management to maximize long-term shareholder value. Perhaps most obviously, standard agency cost theory teaches that whenever managers do not own 100% of the firm’s residual claims their incentives are not perfectly aligned with those of shareholders.140Jensen & Meckling, supra note 18, at 312–13. Concern about this problem, of course, is as old as the business corporation itself. See Adam Smith, The Wealth of Nations 124 (P.F. Collier & Son 1902) (1776) (“The directors of such companies, however, being the managers rather of other people’s money than of their own, it cannot well be expected, that they should watch over it with the same anxious vigilance with which the partners in a private copartnery frequently watch over their own . . . . Negligence and profusion, therefore, must always prevail, more or less, in the management of the affairs of such a company.”). See generally Adolf A. Berle Jr. & Gardiner C. Means, The Modern Corporation and Private Property (1932) (analyzing agency problems generated by the separation of ownership from control in public companies). The literature on such incentive problems is vast, and we will not rehearse it all here. For present purposes we concentrate on the main incentive problems that result in failure to engage in forms of CSR that would benefit shareholders.

Perhaps the primary incentive problem related to ESV is corporate “short-termism,” in which management focuses myopically on short-run profitability at the expense of long-term shareholder value.141See, e.g., The Global Compact, supra note 44, at 5 (“The use of longer time horizons in investment is an important condition to better capture value creation mechanisms linked to ESG factors.”). A key premise of the standard short-termism argument is that the firm’s stock price does not fully reflect what management knows about the value of the firm, for example, because of information asymmetries between managers and investors.142See Jeremy C. Stein, Takeover Threats and Managerial Myopia, 96 J. Pol. Econ. 61, 62 (1988). Consider the following stylized example. Suppose that managers had private information that an expenditure of $80 (for example, additional investment in research and development) would increase expected revenues by $100. However, investors—because they lack managers’ private information—place only 50% probability on revenues increasing by $100 and 50% probability on revenues remaining the same from this investment.143This example draws on the formal model presented in Stein’s article. See generally id. As a result, investors would view the investment as having an NPV of -$30, whereas managers would view the investment as having an NPV of $20. In this fashion, the company’s stockholders might undervalue a change in a company’s operations that would increase long-term shareholder value.

For such market myopia to actually affect corporate decision-making, however, some sort of “transmission mechanism” must exist that induces corporate management to focus on increasing the company’s short-term stock price rather than long-term shareholder value.144Mark J. Roe, Corporate Short-Termism—In the Boardroom and in the Courtroom, 68 Bus. Law. 977, 985 (2013). One potential such mechanism is the corporate takeover market.145Stein, supra note 142, at 63; Martin Lipton, Takeover Bids in the Target’s Boardroom, 35 Bus. Law. 101, 109 (1979). In particular, managers might be concerned that if the market undervalues the long-term value of a particular strategy, a corporate raider might exploit the temporary mispricing in the company’s stock and acquire the company at a price that does not reflect the long-term value of the company, thus deterring managers from undertaking the strategy. In today’s corporate landscape, however, a more common version of this concern involves hedge fund activists who take only a minority stake in a target and then agitate for operational or financial changes that might increase the company’s share price even if the changes undermine long-term shareholder value.146Martijn Cremers, Saura Masconale & Simone M. Sepe, Activist Hedge Funds and the Corporation, 94 Wash. U. L. Rev. 261, 270–71 (2016). As with corporate takeovers, even just the threat of such activist interventions might produce managerial myopia more broadly by incentivizing management to pay excessive attention to short-term results for fear of the company becoming a target.147Robert Kuttner, The Truth About Corporate Raiders, New Republic, Jan. 20, 1986, at 14, 17; cf. Stein, supra note 142, at 63 (“In [takeover] cases, managers who boost their stock prices by inflating earnings may be attempting to act in the interests of stockholders by preventing them from being unfairly ‘ripped off’ by raiders.”). Even more directly, modern executive compensation packages generally make managers themselves short-term stockholders, and there is some evidence that vesting equity induces CEOs to cut back on long-term corporate investments148Alex Edmans, Vivian W. Fang & Katharina A. Lewellen, Equity Vesting and Investment, 30 Rev. Fin. Stud. 2229, 2231 (2017). and to engage in stock repurchases and corporate acquisitions that impair long-term shareholder returns.149Alex Edmans, Vivian W. Fang & Allen H. Huang, The Long-Term Consequences of Short-Term Incentives, 60 J. Acct. Rsch. 1007 (2022). Corroborating the hypothesis that short-termism might inhibit both firm performance and CSR investments is evidence that both firm performance and investments in stakeholder relationships increase as a result of reforms that improve executives’ long-term incentives.150Caroline Flammer & Pratima Bansal, Does a Long-Term Orientation Create Value?: Evidence from a Regression Discontinuity, 38 Strategic Mgmt. J. 1827, 1827 (2017). The extent of managerial short-termism remains controversial,151For a skeptical view, see generally Mark J. Roe, Stock Market Short-Termism’s Impact, 167 U. Pa. L. Rev. 71 (2018). Similarly, for a positive view of hedge fund activism, in terms of long-term shareholder value effects, see Lucian A. Bebchuk, Alon Brav & Wei Jiang, The Long-Term Effects of Hedge Funds Activism, 115 Colum. L. Rev. 1085, 1121–35 (2015). but it provides a coherent conceptual account for why corporate managers might sometimes fail to engage in CSR that would ultimately increase long-term shareholder value.

Other kinds of agency problems can also inhibit CSR under the ESV approach. For instance, managers might engage in empire building or otherwise overinvest in ways that harm long-term shareholder value. For firms that operate in high-negative-externality industries—fossil fuel production, say—such overinvestment can harm other interests in society as well. Alternatively, disloyal managers might claim to sacrifice short-term profitability to further stakeholder interests in the name of long-term value creation when in fact they are engaged in a form of self-dealing.

To summarize, management pursuit of ESV is neither hopeless nor a sure thing. We can expect corporate managers to be able to gather and analyze a substantial amount of the information needed to engage in CSR under the ESV approach and to have considerable incentives to do so, but their information and incentives will not be perfect.

2.  Shareholder Social Preferences

Consider now the extent to which changing the corporate objective from long-term shareholder value under ESV to shareholder welfare under the SSP approach is likely to make corporate conduct more socially responsible. For this reform to achieve its goal of increased corporate social responsibility, corporate managers need both information about their shareholders’ social preferences and incentives to act on that information.

i.  Sorting of Shareholders

A key premise of the SSP approach is that shareholders have social preferences that make them willing, in aggregate, to sacrifice shareholder value in order for the corporation to act more in line with their values. But as an initial matter, will socially minded investors actually be willing to hold the stock of companies whose operations raise the greatest social concerns? So far, we have maintained the simplifying assumption that all shareholders are perfectly diversified. In practice, however, shareholders’ incentives to hold the shares of a particular issuer will in fact depend on their social preferences. This is because shareholders’ social preferences are, at least in important part, associative. By associative we mean that shareholders prefer not to own shares in (or otherwise be associated with) companies whose business practices they find morally objectionable. One source of evidence for this stems from the portfolios of ESG mutual funds that are marketed to appeal to such investors, which are tilted towards companies with high ESG scores.152Quinn Curtis, Jill Fisch & Adriana Z. Robertson, Do ESG Mutual Funds Deliver on Their Promises?, 120 Mich. L. Rev. 393, 424 (2021). In turn, mutual funds marketed as socially responsible are disproportionately held by more prosocial investors.153Arno Riedl & Paul Smeets, Why Do Investors Hold Socially Responsible Mutual Funds?, 72 J. Fin. 2505, 2507 (2017). Individuals’ direct holdings of stock exhibit a similar phenomenon. In particular, individuals who vote in favor of shareholder proposals pressuring the company to act more responsibly are more likely to hold renewable energy firms and less likely to hold fossil-fuel producers. Jonathon Zytnick, Do Mutual Funds Represent Individual Investors? 39 (NYU L. & Econ., Research Paper No. 21-04, 2022), https://papers.ssrn.com/abstract=3803690 [https://perma.cc/X2MQ-ZHMD] (“[I]ndividuals who vote in favor of SRI proposals are more likely to own renewable energy firms and less likely to own fossil fuel producers.”). The result of such shareholder sorting is to further reduce the importance of shareholder social preferences in the shareholder welfare objective function for the very corporations for which there is the most at stake in terms of CSR. The shareholders that hold companies that raise the greatest social concerns will be systematically the investors least concerned about those social issues.154See Ľuboš Pástor, Robert F. Stambaugh & Lucian A. Taylor, Sustainable Investing in Equilibrium, 142 J. Fin. Econ. 550, 553–57 (2021) (developing a model of investing in an economy in which investors differ in their degree of concern about corporate social behavior and showing that, in equilibrium, dirty firms are disproportionately held by investors least concerned about corporate social behavior).

Hart and Zingales, in proposing the SSP approach, in contrast adopt a very different assumption about the form of investors’ social preferences and how they manifest in behavior. They assume that shareholders care about corporate behavior only at the point they are asked to make some decision about it—like voting on a shareholder proposal—and not before or after such a shareholder decision is made.155Hart & Zingales, supra note 71, at 253. They adopt the same approach in their later work on shareholder social preferences. See Broccardo et al., supra note 78, at 3103. Under their view, environmentalists would have no qualms about owning shares in a coal-mining company. Their social preferences would manifest only if they were asked to decide on some specific operational matter that would implicate their environmentalist views. If shareholders were asked to vote on whether the company should adopt a more environmentally responsible mining technique, say, that would lower shareholder returns to some extent, environmentalist shareholders might vote yes, depending on the weight they put on their environmentalist views and the extent of the lower shareholder return entailed. But under Hart and Zingales’s view they would not hesitate to invest in the first place, even if there were no prospect for them to influence the firm’s environmental practices. Hart and Zingales thus propose an invest and engage model of socially responsible investing. But if shareholders’ social preferences are strongly associative, as existing evidence suggests, then this model would work only for companies with operations that are already relatively socially responsible, substantially undercutting the potential of SSP to improve corporate conduct.156To be sure, it could be that the current practice of associative avoidance rather than invest and engage is a function of current corporate governance institutions oriented around shareholder value. Although we are skeptical, it is possible that moving to the SSP regime could cause shareholders to change their sorting behavior and adopt an invest and engage model of socially responsible investment. But such a shift would require that the SSP approach make a substantial difference in corporate behavior, and in what follows we provide further reasons to believe that it would not. See infra notes 159–175 and accompanying text.

ii.  Information

In order for the shift to shareholder welfare as the corporate objective to affect corporate behavior in the intended way, managers must have information about their shareholders’ aggregate social preferences. Relevant preference information would include shareholders’ willingness to pay, in terms of reduced shareholder returns, to further various social concerns as well as how shareholders view trade-offs among competing social concerns. A natural way to gather such information would be for corporate management to poll their shareholders.157See Hart & Zingales, supra note 71; Alex Edmans & Tom Gosling, How To Give Shareholders a Say in Corporate Social Responsibility, Wall St. J. (Dec. 6, 2020, 11:00 AM ET), https://www.wsj.com/articles/how-to-give-shareholders-a-say-in-corporate-social-responsibility-116072

70401 [https://perma.cc/5U8E-4BKW] (arguing in favor of periodic shareholder votes on “corporate purpose” as a way for management to elicit information about shareholders’ social preferences); Jill E. Fisch, Purpose Proposals, 1 U. Chi. Bus. L. Rev. 113, 128–55 (2022) (analyzing purpose proposals).

One version of this would be for management to poll shareholders for their views on concrete corporate operational matters that implicate various social concerns. As a preliminary matter, however, note that diversified shareholders generally lack the information and expertise needed to understand the trade-offs available between firm value and social concerns—this is the core economic logic of centralized management. Put simply individual shareholders are unlikely to know what corporate decisions would maximize their utility.

Consider, for example, the shareholders of a social-media company. Many of these shareholders might share a belief that the corporation should protect the privacy and data of its users, but they likely have little knowledge of the different corporate practices that could advance those interests and the trade-offs they would entail. In principle, shareholders could, with the help of management, inform themselves of the relevant options and their associated costs, but doing so would entail costs that would likely deter diversified shareholders from doing so.158Skepticism regarding whether shareholders are well-positioned to evaluate specific corporate policies also appears in the SEC’s policy of excluding 14a-8 proposals that seek “to ‘micro-manage’ the company by probing too deeply into matters of a complex nature upon which shareholders, as a group, would not be in a position to make an informed judgment.” Amendments to Rules on Shareholder Proposals, Exchange Act Release No. 34-40018, 63 Fed. Reg. 102, at 29109 (May 28, 1998).

Consider then instead the possibility that management might learn information just about the content and strength of shareholders’ social preferences rather than shareholders’ views about specific operational decisions. Even at this raw preference level, however, we are skeptical that shareholders have clear preferences in any meaningful sense about the relevant trade-offs, much less that management could realistically learn much about them. For example, consider again a social-media company. Another major social concern about social media is its role in the spread of disinformation. Suppose you, dear reader, were a shareholder of a social-media company that had been plagued by such problems in the past. How much return would you be willing to sacrifice in order to reduce this problem? If you are like us, you are having trouble even coming up with a coherent metric for expressing such a preference. Are you willing to sacrifice fifty basis points in return for a reduction of one . . . disinformation unit?

Put another way, shareholder voting provides information about the stated preferences of shareholders but not necessarily their revealed preferences. As a result, a risk exists that asking what any given shareholder prefers in terms of social issues and investment returns might result in the shareholder expressing a preference that is inconsistent with the policy the shareholder would adopt if forced to pay directly for the policy adoption.159Economists are traditionally skeptical of using stated preference methods for eliciting individuals’ valuations of public goods and the like as a guide for welfare analysis. After surveying the empirical literature documenting biases and inconsistencies in responses to surveys eliciting individuals’ valuations of various environmental amenities, Peter Diamond and Jerry Hausman conclude that the problems with such stated preference methods:

[C]ome from an absence of preferences, not a flaw in survey methodology. That is, we do not think that people generally hold views about individual environmental sites (many of which they have never heard of); or that, within the confines of the time available for survey instruments, people will focus successfully on the identification of preferences, to the exclusion of other bases for answering survey questions. This absence of preferences shows up as inconsistency in responses across surveys and implies that the survey responses are not satisfactory bases for policy.

Peter A. Diamond & Jerry A. Hausman, Contingent Valuation: Is Some Number Better than No Number?, 8 J. Econ. Persps. 45, 63 (1994).
As well, we might question whether preference elicitation is in the wheelhouse of corporate managers.

These informational challenges facing SSP are not much diminished when we consider intermediation by institutional investors. Hart and Zingales propose that such intermediaries might provide a means of lowering the cognitive load on diversified investors of expressing their social preferences over corporate conduct.160Hart & Zingales, supra note 71. Prosocial investors could simply invest in a prosocial mutual fund that will vote its portfolio company shares in order to advance the investors’ social preferences. But this essentially just moves the information problem down one level: How can the fund’s manager learn about the social preferences of its investors in order to relay that information to corporate managers?161In response to this challenge, one could, of course, require institutional investors to solicit the views of their investors and vote accordingly. See Jill E. Fisch & Jeff Schwartz, Corporate Democracy and the Intermediary Voting Dilemma, 102 Tex. L. Rev. 1, 48 (2023) (proposing “a system by which fund managers ascertain the preferences of their beneficiaries and incorporate those preferences into their voting and engagement practices”). Despite its appeal, such an approach would hardly be a mechanism for implementing SSP for several reasons. First, this form of polling would have to overcome the problem of investor passivity in corporate voting. See Alon Brav, Matthew Cain & Jonathon Zytnick, Retail Shareholder Participation in the Proxy Process: Monitoring, Engagement and Voting, 144 J. Fin. Econ. 492, 500 (2022) (finding only 11% of retail accounts cast votes at annual shareholder meetings). More importantly, soliciting investors’ general preferences on social issues would similarly suffer from its inability to capture investors’ revealed preferences on the concrete trade-offs implicated by specific voting proposals. Indeed, even advocates of this approach acknowledge the continuing need for institutional investors to engage in informed intermediation given that whatever preferences are expressed through such polling are likely to be “incomplete, inconsistent, or uninformed.” Fisch & Schwartz, supra, at 9. As such, there could be no assurance that the votes cast by institutional investors would, in fact, reflect the true preferences of a company’s beneficial owners.

One possibility, suggested by Hart and Zingales, is that investors can “vote with their feet” by sorting into funds that have a track record of voting that investors find attractive.162Hart & Zingales, supra note 71, at 265. Indeed, Michal Barzuza, Quinn Curtis, and David Webber argue that index-fund providers have become increasingly vocal about their voting records on ESG issues in order to compete for millennial investors, who they argue place a significant premium on social issues.163See Michal Barzuza, Quinn Curtis & David H. Webber, Shareholder Value(s): Index Fund ESG Activism and the New Millennial Corporate Governance, 93 S. Cal. L. Rev. 1243, 1265–68 (2020).

But empirical evidence provides little support for the idea that investors sort into mutual funds based on their voting policies. For instance, using a dataset that contains the voting records of both individual investors and the mutual funds in which they invest, Jonathon Zytnick examines whether mutual funds vote on CSR-related matters in the same way that their investors vote on CSR-related matters when these investors cast ballots as shareholders.164Zytnick, supra note 153, at 27–36. Overall, he finds little overlap between investor preferences and fund voting, especially within index funds.165Id. at 29. One exception is with respect to ESG funds, which typically vote in favor of CSR-related initiatives, which is consistent with how their investors cast ballots as individual shareholders. But note that ESG funds typically focus on screening out firms with poor ESG track records, reflecting our view that investors’ social preferences are to a large extent associational. Id. at 29–31. Zytnick attributes the overall lack of sorting to rational inattention: as in political voting, investors rationally choose not to investigate how an intermediary votes due to the small likelihood that their investment will cause the intermediary’s votes to be pivotal.166Id. at 19.

Hart and Zingales argue that the lack of investor sorting is due to current corporate governance rules that limit the scope of shareholder voting on CSR.167Hart & Zingales, supra note 71, at 264. State corporate law, for example, gives the board but not shareholders the legal authority to manage the business and affairs of the corporation. This norm prevents shareholders from restricting the board’s substantive decision-making authority by enacting bylaws that direct particular substantive outcomes in terms of CSR. See CA, Inc. v. AFSCME Emps. Pension Plan, 953 A.2d 227, 234–35 (Del. 2008). In turn, Rule 14a-8 of the federal proxy rules, which gives shareholders the right to put certain shareholder proposals on management’s proxy for the annual shareholder meeting, allows management to exclude proposals that are “not a proper subject for action by shareholders under the laws of the jurisdiction of the company’s organization.” 17 C.F.R. § 240.14a-8(h)(3)(i) (2022). However, even in the absence of such limitations, we question whether sorting among funds based on how they vote on social issues would provide meaningful information to managers about their shareholders’ social preferences. First, as we argued above, we doubt that investors have sufficiently well-formed preferences about corporate conduct such that it is even possible for sorting to convey information to corporate managers about those preferences. Second, it would remain prohibitively costly for shareholders to evaluate the stated policies of asset managers. It is not as simple as environmentally minded shareholders buying a “green” mutual fund. As we have emphasized, shareholders’ social preferences are heterogeneous, both in terms of their strength relative to wealth in their utility function and in terms of their content. Individual investors will often differ in how they evaluate the trade-offs entailed when a company implements specific CSR-related policies.

Consider, for example, a fund dedicated to carbon reduction. Across the range of policy interventions a company might take to reduce its carbon footprint, how will investors know which ones a Reduce Carbon Fund will pursue, or how it will evaluate the inevitable trade-offs implicated by each course of action? While some investors may adopt a hell-or-high-water (“hah”) approach to carbon reduction, others may condition their support on evidence that the intervention will enhance long-term shareholder value. These problems are further compounded in cases in which a corporate decision involves a trade-off between competing social values and not just between a single social issue and investment returns. Many shareholders, for example, might have concerns about the implications of a given carbon reduction policy proposal on other stakeholders, such as workers or communities who may be adversely impacted by it.168Indeed, BlackRock, which is the largest asset manager in the United States, announced a new program in January 2022 called “Voting Choice” whereby it will allow its clients to choose how to vote the portfolio securities of certain BlackRock funds managed on their behalf. Shareholder Rights Directive II — Engagement Policy, BlackRock (2022), https://www.blackrock.com/corporate/
literature/publication/blk-shareholder-rights-directiveii-engagement-policy-2022.pdf [https://perma.cc/
Y3N5-8JN3]. While the initial program includes only institutional clients, the firm has announced that it is “committed to a future where every investor—even individual investors—can have the option to participate in the proxy voting process if they choose.” Fink, supra note 101. On the one hand, these changes might be thought of as facilitating the SSP approach by enabling shareholders who invest through intermediaries to express their views on social issues. But we suspect that this emerging devolution of voting responsibility to beneficial owners reflects both the difficulties asset managers face in determining their investors’ preferences and the intractability of the conflicts among shareholders in their social preferences. These changes enable asset managers to sidestep these issues and push down the costs of becoming informed on the issues being voted on to their underlying investors, who lack incentives to bear them, ultimately undermining the feasibility of the SSP approach.

A final problem with using shareholder voting and similar mechanisms to convey social preference information to corporate management is that shareholders will express their overall preferences about corporate policy, not just the part concerning shareholder value and their social preferences. For diversified shareholders, those overall preferences would include the portfolio effects that PVM—and not SSP—envisions incorporating into the corporate objective. As a result, attempts to implement the SSP approach, to the extent they are successful in tilting corporate decisions toward what shareholders want, will in practice blur into pursuit of the PVM objective including the anticompetitive aspects of it that are socially destructive.

iii.  Incentives

As we argued above, shareholder value is a much more important component of shareholder welfare than shareholder social preferences, given heterogeneity and conflicts among shareholders regarding the relevant welfare trade-offs and sorting based on associative preferences. Our analysis also revealed that management has much better information about long-term shareholder value than it has about shareholders’ social preferences. In such a setting—with one far more important component of the objective function for which information is readily available and one far less important component for which information is not available—the best scheme for incentivizing corporate management to pursue shareholder welfare under the SSP approach focuses management attention squarely on the important and measurable component, long-term shareholder value, and thus is essentially identical to the ESV approach.

Our argument builds on insights from “multitask principal-agent problems” from contract theory.169Bengt Holmstrom & Paul Milgrom, Multitask Principal–Agent Analyses: Incentive Contracts, Asset Ownership, and Job Design, 7 J.L. Econ. & Org. 24, 25 (1991). These models entail a principal who hires an agent to perform several tasks or, similarly, a single task with multiple dimensions to it. A common problem in such an environment arises when performance on one dimension of the job is easily measurable while performance on another dimension is difficult to measure. Teacher performance is a classic example. Standardized tests can measure one dimension of teacher performance, but other aspects—promoting creativity or communication skills—are much harder to measure. In such a setting, the agent decides how to allocate effort across the dimensions of the job, and an increase in incentives on the more easily measurable dimension of their performance will result in the agent reallocating their effort toward that dimension and away from the others.

In a pathbreaking article working through the implications of such a setting for contract design, Bengt Holmstrom and Paul Milgrom argued that the optimal contract might entail very low-powered incentives, like a fixed wage, in order to avoid distorting the agent’s effort too much in the direction of the more easily measurable dimension of the job.170Id. at 35–38. In the application to teachers, the idea is that paying teachers based on a fixed salary would result in better overall teacher performance than paying them based on the performance of their students on standardized tests since the more balanced allocation of teacher effort across the different dimensions of their job that would result—based on teachers’ intrinsic motivations—is more important than the fall in overall effort from giving up on high-powered extrinsic incentives on the measurable aspect of their performance.

In our setting, a low-powered incentive contract in the spirit of Holmstrom and Milgrom’s analysis would entail giving up on providing managers high-powered incentives to maximize shareholder value in order to induce them to put some effort into measuring and furthering shareholders’ social preferences. For example, managers could be paid like bureaucrats, with fixed salaries and no equity-based component to their pay. But this is not the optimal contract here, for two reasons.

First, as we have explained, long-term shareholder value is a more important component of shareholder welfare than is shareholder social preferences—by far—and, in addition, managers have much better information about how to maximize shareholder value than about how to satisfy shareholders’ social preferences. As a result, managerial effort to maximize shareholder value is generally much more productive, in shareholder welfare terms, than is managerial effort to further shareholders’ social preferences. Consider, then, how shareholders would ideally want managers to allocate their finite time and attention across those two tasks. For the sake of argument, suppose that management were to focus exclusively on maximizing shareholder value and ignored shareholders’ social preferences. From this benchmark, would shareholders’ welfare increase if management were to divert some of its attention to figuring out how best to further shareholders’ social preferences? We think not. The resulting fall in shareholder value would matter more to shareholder welfare than whatever small improvement management could achieve in better aligning firm policy with shareholders’ social preferences.

Second, suppose we are wrong about that, and in fact shareholders would ideally want management to devote at least some attention to furthering shareholders’ social preferences. That alone is not sufficient for the optimal incentive contract for management to be one that avoids high-powered incentives to maximize firm value. The optimal design of incentives depends not only on the relative productivity of management’s efforts on the two tasks but also on management’s intrinsic motivation to pursue the tasks as well as on the availability of good incentive instruments to motivate managerial effort on each of the tasks.

In the application of the Holmstrom and Milgrom multitask model to the problem of incentivizing teachers, a fixed wage contract results in teachers’ effort being driven by their intrinsic motivation to help students learn. In the educational context, it seems plausible that teachers have substantial intrinsic motivation—presumably many teachers enter the profession not because the pay is high (it is not) but rather because they like teaching and care about students. As a result of their intrinsic motivations, the fixed wage contract for teachers results in substantial effort across both the measurable and nonmeasurable dimensions of their performance.

But in the corporate context, we think intrinsic motivations play a much smaller role relative to extrinsic motivations. As a result, giving up on extrinsic incentives would result in a substantial fall in managerial effort on maximizing firm value, and for little benefit; it is hard to see why corporate managers would have much intrinsic motivation to figure out shareholders’ social preferences and seek to further them.

In terms of the availability of incentive instruments, the key issue is whether there are good proxies for the agent’s performance to base their compensation on. When an agent is paid on the basis of some performance measure, they will have incentives to increase the performance measure, which might not produce the desired results. The basic analytic point here is captured evocatively in the title of a classic article in the management literature: On the Folly of Rewarding A, While Hoping for B.171Steven Kerr, On the Folly of Rewarding A, While Hoping for B, 18 Acad. Mgmt. J. 769 (1975). In the teacher context, there might not be great proxies even for the relatively measurable aspects of the job. Consider the practice of paying teachers based on their students’ test scores. The hope is that doing so will motivate teachers to teach better. But following the introduction of incentive pay based on test scores for teachers in Atlanta, ten teachers and administrators were caught helping students cheat on the test to inflate their scores.172Annie Murphy Paul, Atlanta Teachers Were Offered Bonuses for High Test Scores. Of Course They Cheated., Wash. Post (Apr. 16, 2015, 12:43 PM EDT), https://www.washingtonpost.com/post
everything/wp/2015/04/16/atlanta-teachers-were-offered-bonuses-for-high-test-scores-of-course-they-cheated [https://perma.cc/E5P5-PBEP].
Put simply: you get what you pay for.

The implication for the optimal design of incentives is that the fall in effort on the measurable dimension of performance from switching from a high-powered incentive scheme to low-powered incentives depends on how well the former dimension of performance can in fact be measured.173See George P. Baker, Incentive Contracts and Performance Measurement, 100 J. Pol. Econ. 598, 599 (1992) (“[T]o the extent that the performance measure does not respond to the agent’s actions in the same way that the principal’s objective responds to these actions, the firm will reduce the sensitivity of the incentive contract to the performance measure.”). In teaching, test scores are a potentially problematic measure even of the aspects of teacher performance they purport to measure, as the cheating scandal illustrates in extreme form. This measurement problem then reduces the benefit, in terms of student learning, of paying teachers based on the proxy. In contrast, in the corporate context, there are excellent performance measures available for shareholder value. The shareholder value component of shareholder welfare is ultimately revealed over time as the firm’s cash flows are realized. Executive compensation plans make use of that fact by employing equity-based pay and explicit bonus schemes tied to accounting measures of earnings to generate incentives to maximize shareholder value. We believe that equity-based pay can provide substantial alignment between management’s incentives and shareholder value. Giving up on those incentives would therefore result in a substantial loss in shareholder value.

Finally, we do not believe it is optimal to add explicit incentives for managers to further shareholders’ social preferences. The shareholder social preferences component of shareholder welfare is much harder to measure than shareholder value and remains largely hidden. Some crude proxy for shareholders’ social preferences, based on surveys of shareholders or the like, would have to be constructed to use as a performance measure in management’s compensation scheme. But the measurement challenges here reduce the productivity, from a shareholder welfare perspective, of trying to provide extrinsic incentives to management to take into consideration shareholders’ social preferences.

In sum, the optimal incentive scheme under the SSP view focuses squarely on shareholder value, so that the SSP approach would do little to improve corporate behavior relative to the ESV baseline. One response might be that there is no downside to changing the corporate objective to shareholder welfare under SSP and possible upside. If we are right, the argument goes, that the optimal incentive scheme would remain unchanged, then boards charged with pursuing shareholder welfare under SSP will ensure that management has incentives to stay focused on shareholder value. But it could be, the argument continues, that for some firms, the information and incentive problems we have identified with seeking to further shareholders’ social preferences are less severe. For those firms, changing the corporate objective to shareholder welfare under SSP could result in more socially responsible corporate behavior. But in our view, such a change to the legal and business norm about corporate purpose would inevitably result in substantial efforts by many corporate boards to induce the company’s senior managers to incorporate shareholders’ social preferences into their decision-making even when doing so in fact lowers shareholder (and social) welfare by distracting management from shareholder value.

3.  Portfolio Value Maximization

Evaluating the feasibility of PVM as an alternative corporate objective requires assessing whether corporate managers might have the information and incentives needed to incorporate the effects of the firm’s decisions on the value of their shareholders’ portfolios into their decision-making process, above and beyond how those decisions affect the long-term value of the corporation. We show here that there are good reasons to think they will not.

i.  Information

A first type of information managers would need under PVM is on the composition of the portfolios held by the company’s shareholders. A company’s shareholders are likely to vary widely in the investment portfolios that they hold. Indeed, the large number of investment products offered as mutual funds reflects the strong demand for a broad range of investment portfolios with varying investment objectives. As of November 2023, Morningstar lists over 1,800 investment funds as providing exposure to “U.S. Equity” and nearly 1,100 investment funds as providing exposure to “International Equity.”174For the list of U.S. Equity funds, see U.S. Equity Funds, Morningstar, https://www.morning

star.com/us-equity-funds [https://perma.cc/24N7-M6WJ]. For the list of International Equity funds, see International Equity Funds, Morningstar, https://www.morningstar.com/international-equity-funds [https://perma.cc/TZ5P-XGCT].
Moreover, the portfolios of these funds reflect a broad range of investment theses, such as funds focused on growth firms, small-capitalization firms, low-volatility firms, dividend-paying firms, or firms operating in particular regions or sectors. Note as well that it is not enough for managers to determine what institutional investors hold the company’s shares. Institutional investors serve as intermediaries for the underlying individuals on whose behalf they ultimately hold the company’s shares. In turn it is those individual investors’ portfolios that form the ultimate aggregate portfolio the company’s managers should be trying to maximize.

To keep things simple, however, suppose corporate management assumed that the company’s shareholders are fully diversified so that the PVM objective is just the value of the market portfolio. This simplifying assumption stacks the deck in favor of the feasibility of PVM, so if PVM is not reasonably feasible under this assumption, then it certainly is not feasible in the real world.

A second type of information a corporate manager would need to pursue PVM is on the expected cash flows that alternative decisions would generate, not only for the company itself but also for other securities in shareholders’ portfolios, which again for now we take to be the market portfolio. These expected cash flows to the company and to other securities in the market portfolio are the ’s and ’s, respectively, in the numerators of the terms in the PVM version of the expression for the NPV of a project in equation (2) above.

In general, corporate managers will have much better information about the cash flows to the company (the ’s) than they will about the portfolio externality cash flows (the ’s). The cash flows to the company are ultimately directly observable and of course directly implicate the business of the company, on which managers are hired to be experts. Externalities, in contrast, involve other businesses that the firm’s managers will have much less information about. The information challenges posed by technological externalities are particularly acute. It is not clear how a firm’s managers would be able to divine the extent to which pollution emitted by the company, say, would reduce the value of other public companies, which include a diverse array of sectors and industries.175To be sure, there might be some specific technological externalities for which these information problems are less substantial. Most notably, there are aspects of the climate change policy problem that make it more amenable to institutional investors and managers having the requisite information. A ton of CO2 emitted in the atmosphere results in the same marginal social costs regardless of where or how it is emitted, since each such ton contributes the same global stock of greenhouse gases in the atmosphere that in turn causes climate change. Accordingly, institutional investors could collaborate with government and other actors to analyze the portfolio effects of climate change, as the UNEP FI has attempted to do. See United Nations Env’t Programme Fin. Initiative, supra note 117, at 38–49. Yet even this setting, in which one can plausibly model the portfolio effects of producing a unit of an externality, ultimately illustrates the limitations of the PVM approach. As we have already noted, the offsetting positive effects of climate change for many publicly traded companies, along with the use of discount rates far above the social discount rate and the geographic mismatch between the market portfolio and the economic costs of climate change, means that the net physical costs of climate change on the market portfolio are likely to be de minimis. See supra notes 116–133 and accompanying text. In contrast, pecuniary externalities primarily affect the company’s competitors, about which firm managers are likely to have substantial information.

Nor are institutional investors likely to be in a meaningfully better position to provide this information to managers. Acquiring information about the ’s of a portfolio company would require a level of firm-specific engagement likely to be far more complex than acquiring information only about the ’s of the company by virtue of the diffuse ways a company’s operations can affect firms in the market portfolio. Yet even when it comes to firm-specific engagement on increasing a company’s ’s, both active asset managers and index-fund providers have strong incentives to refrain from active engagement.176For active managers, any action that increases the value of a portfolio company will be shared by all active managers holding a position in the company; therefore, the initiating manager will suffer a decline in relative performance to the other managers who will similarly benefit from the increase in the company’s value without having to incur the costs of engagement. See Ronald J. Gilson & Jeffrey N. Gordon, The Agency Costs of Agency Capitalism: Activist Investors and the Revaluation of Governance Rights, 113 Colum. L. Rev. 863, 891–92 (2013). Likewise, index providers compete for assets under management on the basis of their low fees, making the costs associated with such firm-specific engagement incompatible with their business model. See id. Rather, both types of institutional investors adopt a stance of “rational reticence”177Id. at 867, 889. in which they weigh in on a company’s operations only after an activist hedge fund—which has incentives to investigate how a company might increase its cash flows due to its concentrated investment position—proposes an intervention. Yet by the same token, the fact that an activist is undiversified also means it has little reason to invest in exploring how to reduce the ’s of a company. Indeed, to the extent an activist surfaces information on a company’s technological externalities, it will most likely relate to how they adversely affect the company’s cash flows—a point to which we return in Part V. As a result, managers cannot count on institutional investors to solve the critical information challenge posed by PVM.178Due to this challenge, Jeffrey Gordon suggests that, in the context of financial stability risk, institutional investors “ought to devote more firm-specific (and sector-specific) attention to financial firms precisely because (i) they cannot rely on some of the standard intermediaries and (ii) a single-firm failure can present a systemic threat.” Gordon, supra note 84, at 660. However, even assuming systemic risk of this sort was confined to preventing the failure of, say, any of the thirty firms listed by the Financial Stability Board as a Global Systemically Important Bank, see Fin. Stability Bd., 2022 List of Global Systemically Important Banks (G-SIBs) 3 (2022), https://www.fsb.org/2022/11/2022-list-of-global-systemically-important-banks-g-sibs [https://perma.cc/UF8J-7Q9T], we question whether active managers and indexers would view active engagement across even these thirty firms as cost justified, given their strong incentives for governance passivity. See Gilson & Gordon, supra note 176, at 891–92. More importantly, the 2023 banking crisis is a stark reminder that efforts to contain financial stability risk would require a far greater expenditure of resources given the interconnectedness of financial institutions. The crisis represents precisely the type of nondiversifiable financial stability risk at the heart of PVM; yet it was initiated by the failure of just three regional banks (Silicon Valley Bank, Silvergate Bank, and Signature Bank). As of December 31, 2022, the Federal Reserve listed 2,214 banks on its list of “large commercial banks” operating in the United States. Large Commercial Banks, Fed. Rsrv. (Dec. 31, 2022), https://www.federalreserve.gov/releases/lbr/20221231 [https://perma.cc/MZR9-XJ9R]. In addition to firm-specific engagement, Gordon also suggests institutional investors could adopt portfolio-wide policies that favor more specific disclosures regarding a company’s exposure to areas of systemic risk, such as through supporting private and quasi-regulatory efforts to provide more uniform disclosure standards on climate change risk. Gordon, supra note 84, at 661. Even here, however, the goal would be to facilitate better pricing of a company’s securities to reflect a company’s exposure to systemic risk. Yet to the extent markets can better price a firm’s exposure to a particular type of systemic risk, this simply ensures investors will be compensated for bearing this form of nondiversifiable risk.

ii.  Incentives

Consider now the implications of the foregoing analysis for the incentives that firm managers have to pursue the PVM objective. The long-term value of the firm’s own shares and the pecuniary portfolio externalities produced by the firm are far more important components of the PVM objective function than the technological portfolio externalities produced by the firm. One reason for this is that there exist social institutions, such as environmental regulation, designed to internalize technological externalities of corporate activity. While these institutions are certainly imperfect, they do substantially limit technological externalities. Another reason is that only a fraction of corporate technological externalities actually falls on other companies’ securities, as we explained above. As a result, when managers are considering investing in a new project, typically the primary effect it has on investors’ portfolios is through its implications for the company’s own value. As well, pecuniary externalities are likely to be far more important to its shareholders than technological externalities for the reasons discussed above. Note that the ordering of these three components of the PVM objective function in terms of their importance to investors mirrors their ordering in terms of the information available to managers.

Incentivizing firm managers to incorporate technological externalities into their decision-making under the PVM approach thus poses a similar problem to that of incentivizing them to consider shareholder social preferences under the SSP approach. The most productive use of managers’ scarce time and attention, in terms of improving the PVM objective function, is in working to increase the cash flows to the firm’s own shares and to competing public companies. As a result, we think it likely that diversified shareholders would want managers to focus their limited time and attention on those outcomes. Diverting their attention to addressing technological portfolio externalities would likely be counterproductive for the value of shareholders’ portfolios, given their relatively small role in the PVM objective function and the relatively limited information firm managers have about them. The optimal incentive contract for managers under the PVM approach would thus focus squarely on the long-term value component of the objective function and put little to no weight on technological externalities.179In the absence of antitrust laws, the optimal incentive contract might also seek to encourage managers to create pecuniary externalities by, for example, colluding with the firm’s competitors. Reforms that aim to induce managers to incorporate portfolio effects into their decision-making are likely counterproductive for both diversified portfolio returns and for social welfare.

These considerations help explain why institutional investors have refrained from pushing managers of high carbon-emitting firms to slash emissions in the name of maximizing the value of other portfolio firms, as one might expect if investors truly wanted firms to adopt a PVM perspective. On the contrary, to the extent investors evaluate the impact of climate change on portfolio value maximization, they typically focus on the implications of climate change for each firm’s long-term value and in particular on transition risks, such as the costs a firm will face as governments seek to rein in carbon emissions and the investment opportunities these efforts will produce.180See, e.g., BlackRock, Climate-Related Risk and the Energy Transition 1 (2023), https://www.blackrock.com/corporate/literature/publication/blk-commentary-climate-risk-and-energy-transition.pdf [https://perma.cc/C69L-ZXJQ] (“While companies in various sectors and geographies may be affected differently by climate change, the energy transition is an investment factor that we expect to be material for many companies and economies around the globe. Within this context, and as stewards of our clients’ assets, we engage companies and encourage them to publish disclosures that help their investors understand how they identify and manage the material risks and opportunities related to climate change and the energy transition.” (endnote omitted)).

Indeed, the work of UNEP FI, which was established to advance methodologies for assessing the impact of climate change on the portfolios of institutional investors, is replete with this perspective. Using an investment portfolio consisting of 30,000 global securities, the report’s headline results indicate that investors in such a portfolio would face a 13.16% risk of loss due to transition risk, but low carbon-technology opportunities offset these costs by providing 10.74% of potential gains. To be sure, the report also estimated the aggregate physical losses to the portfolio arising from climate change to be 2.14%.181United Nations Env’t Programme Fin. Initiative, supra note 117, at 12. Yet even in this regard, the report cited investors as using these methods to engage with companies “to encourage greater climate risk resiliency”—in other words, to ensure companies are looking to maximize firm value in the face of these climate risks.182Id. at 78. Likewise, to the extent shareholder engagement at Big Oil firms has resulted in revised compensation plans to address climate change, the revised plans are uniformly designed to reward management for success in managing transition risk—a broad category of conduct that includes meeting greenhouse gas (“GHG”) emissions targets in anticipation of higher carbon costs as well as pursuing alternative energy technologies.183For instance, in 2021, Chevron approved the addition of an “Energy Transition” performance category to the Chevron Incentive Plan (“CIP”) scorecard in response to investor communications. Chevron Corp., 2022 Proxy Statement (Schedule 14A) 44 (Apr. 7, 2022). According to the company, the “new category will have a 10% weighting, and will measure Chevron’s progress in the areas of GHG management, renewable energy and carbon offsets, and low-carbon technologies.” Id. at 49. In addition to the 10% weight provided to this Energy Transition metric, the CIP determines annual awards based on three other areas: financial results (weighted 35%), capital management (weighted 30%), and operating and safety performance (weighted 25%). Id. at 45.

C.  Devolving Corporate Control to Shareholders

In the prior Section we took as given the current institutional arrangements that give the board of directors control over corporate policy. This model of corporate governance necessarily raises the challenges of how shareholders might convey their preferences to managers (whether to maximize portfolio value or pursue social preferences) as well as how to provide managers with incentives to pursue these preferences. As we have argued, these challenges are difficult—if not impossible—to overcome, so it is hardly surprising that some proponents of shareholder welfarism, from both the SSP and PVM strands, have proposed implementing the shift away from shareholder value maximization toward shareholder welfare maximization by simply giving shareholders much greater direct say in operational matters. This approach is perhaps most associated with two 2022 papers penned by Oliver Hart and Luigi Zingales,184See, e.g., Broccardo et al., supra note 78, at 3101; Oliver Hart & Luigi Zingales, The New Corporate Governance, 1 U. Chi. Bus. L. Rev. 195 (2022). but similar admonitions to provide shareholders with greater voice in corporate governance have long emanated from proponents of PVM.185See, e.g., Hawley & Williams, supra note 84, at 144 (proposing that the governance role of institutional investors should reflect the broader powers of ownership in a corporation, including ‘actively participating in its strategic direction’ ”); Wolf-Georg Ring, Investor Empowerment for Sustainability, 74 Rev. Econ. 21, 21 (2023) (“[F]or investor empowerment as the main tool towards achieving greater sustainability in capital markets” and grounding this “trust in institutional investors . . . in various recent developments both on the supply side and the demand side of financial markets, and also in the increasing tendency of institutional investors to engage in common ownership.”).

We therefore conclude our evaluation of shareholder welfarism by considering the extent to which devolving corporate control to shareholders might improve corporate conduct. Note that, under this implementation mechanism, the distinction between the SSP and PVM forms of shareholder welfarism becomes less significant: in exercising their control rights over a corporation, shareholders would be motivated by their full range of relevant preferences, including with respect to the value of the firm, the value of other securities in their portfolios, and their social preferences. As such, we refer collectively to scholars taking this particular approach to implementing either SSP or PVM as proponents of “shareholder welfarism.”

1.  The Economic Logic of Centralized Control

To begin, we note that adopting a more holistic understanding of shareholder interests, as urged by these proponents of shareholder welfarism, does not change the basic economic logic that originally gave rise to the centralized management of publicly traded corporations. Diversified shareholders generally lack the information and expertise needed to run the firm; this is why, under current institutional arrangements, corporate control is vested in an elected board of directors. Put simply, centralized management lets managers be managers and investors be investors, and that specialization of function has well-understood economic benefits. In our view, devolving operational decisions to shareholders of publicly traded corporations would make little economic sense and would result in worse corporate performance, not just in terms of shareholder value but even in shareholder welfare or social welfare terms.

2.  Determining Which Decisions to Devolve to Shareholders

To be sure, proponents of shareholder welfarism do not propose that all operational decisions be devolved to shareholders, presumably in large part because they recognize the value, indeed practical necessity, of a significant degree of centralization of control over public companies in professional managers. But what then determines which operational decisions are made by shareholders and which by managers? Hart and Zingales argue that, as a conceptual matter, shareholders be given a direct say only with respect to operational issues that implicate a social goal that the company has a comparative advantage in achieving.186Hart & Zingales, supra note 184, at 210. They offer as an example a case from 1984 when DuPont faced a choice between polluting the Ohio River or spending money to avoid doing so.187Id. at 210–11.

But identifying conceptually a class of decisions that should be delegated to shareholders is on its own not enough. One must also specify who decides on a day-to-day basis when a particular corporate decision meets the specified criteria for devolution to shareholders. One possibility is that management decides. We suspect, however, that such an arrangement would result in management rarely bringing matters to a shareholder vote, given the time and expense involved and the fact that shareholders are so poorly equipped to make such decisions. It is not clear why management would have any incentive to bring such votes, and enforcement of a legal obligation for them to do so would presumably entail suits brought by shareholders, in effect making shareholders the key actors in instigating these shareholder votes over corporate operations.

Accordingly, the only plausible approach is to let shareholders initiate such votes, perhaps with management having access to a legal procedure for refusing to bring the vote if it does not meet the specified legal criteria.188This is how Hart and Zingales propose to implement SSP. Id. at 215. This is how the process for putting precatory shareholder proposals on management’s proxy statement for the annual shareholder meeting generally works currently under Rule 14a-8. But consider the incentives of shareholders to initiate such interventions. Standard collective action problems would inhibit diversified individual shareholders from bearing the considerable costs of putting operational issues to a shareholder vote. Similarly, traditional asset managers likely have little incentive to bear the costs of intervening by sponsoring shareholder proposals.189See Gilson & Gordon, supra note 176, at 894.

Consistent with this analysis, existing evidence on precatory shareholder proposals on social issues shows they are proposed largely by what Roberto Tallarita calls “stockholder politics specialists”: policy advocacy organizations like As You Sow, socially responsible investment advisors like Domini Impact Investments, and public and union pension funds.190Roberto Tallarita, Stockholder Politics, 73 Hastings L.J. 1697, 1740–42 (2022). These specialists generally have particular social and political agendas that existing scholarly commentaries characterize as different from the interests of most of the shareholder base.191See, e.g., Susan W. Liebeler, A Proposal to Rescind the Shareholder Proposal Rule, 18 Ga. L. Rev. 425, 439 (1984); Roberta Romano, Public Pension Fund Activism in Corporate Governance Reconsidered, 93 Colum. L. Rev. 795, 807 (1993). It seems likely that these actors often make proposals designed not to push corporate managers to strike a trade-off desired by shareholders between firm value and shareholders’ other preferences (which would be consistent with the view taken by proponents of shareholder welfarism), but rather they make proposals aimed at advancing a particular political agenda. In line with that understanding, only 3.3% of shareholder proposals on social issues from 2010 to 2021 received majority shareholder support.192Tallarita, supra note 190, at 1719. This fraction increased dramatically at the end of the sample period, however, reaching 12.4% in 2019 and 19.2% in 2021. Id. at 1727. Specific categories of social proposals that have begun attracting majority shareholder support at greater rates include proposals on board diversity, climate-related proposals, and proposals on corporate political activity. EY Ctr. for Bd. Matters, Ernst & Young, What Boards Should Know About ESG Developments in the 2021 Proxy Season 3–4 (2021).

We would expect these same actors to be the primary proponents of shareholder proposals under the reforms urged under the shareholder welfarism view that would make shareholder proposals on operational issues binding. The key question is whether empowering these actors to initiate shareholder decisions that override management through binding shareholder resolutions on operational matters is likely, on net, to improve corporate behavior.

3.  The Nature of Shareholder Preferences over Operational Decisions

Consider now how shareholders would vote on proposals pertaining to operational decisions. In an influential article published in the Journal of Political Economy,193Broccardo et al., supra note 78, at 3101. which we will refer to as BHZ, Eleonora Broccardo, Oliver Hart, and Luigi Zingales develop a model of shareholder voting and derive a startling result: in voting over operational decisions that pose trade-offs between firm value and social concerns, diversified shareholders will ignore the implications of the decision for their own investment returns and instead view the decision exactly as a social planner would, making the decision on the basis of the net social benefits to society as a whole.194Id. at 3115. They thus show that, under their assumptions, if a majority of shares are held by investors who are even slightly socially responsible, letting shareholders decide on operational matters achieves the socially optimal outcome. If their model provides a good account of shareholder voting behavior, then devolving operational decision-making to shareholders would have enormous potential for improving corporate conduct. Specialist actors with various views on social issues implicated by corporate conduct could tee up a range of binding resolutions for shareholders to vote on, and shareholders would pass them if and only if they improve social welfare.

But BHZ’s stark result depends on a set of critical assumptions and seems to us implausible in practice. BHZ models investors’ utility from owning a stock as having two components: one stemming from their investment returns from the stock and an altruistic component stemming from how the company’s operations affect society.195Id. at 3113–14. BHZ assumes that, because any individual stock would make up a de minimis fraction of a perfectly diversified investor’s portfolio, such an investor would have no (or de minimis) concern about the effect of an operational decision on their own investment returns.196Id. at 3115. Of course, the assumption of perfectly diversified, atomistic shareholders is inconsistent with how many shares are held, but we put that objection to the side. On the other hand, BHZ assumes that diversification has no effect on the strength of an investor’s ethical concerns about the company’s behavior.197Mathematically, BHZ denotes the number of firms in a diversified portfolio as 𝑟 and uses a utility function in which the investment returns term is multiplied by 1/𝑟, but the social preferences term is not multiplied by 1/𝑟. As a result, in the limit as 𝑟 becomes very large, the investment-returns term goes to zero so that all that is left is the term representing the investor’s social preferences. Id. at 3115. This asymmetry in their treatment of the effects of diversification is the key behind their result that each investor would vote on operational decisions just like a social planner would.

A natural alternative model of investor psychology is from earlier work by Hart and Zingales in which they assumed that the level of responsibility that shareholders feel for corporate externalities scales with their holdings in the firm.198Hart & Zingales, supra note 71, at 253 n.14 (“We suppose that a consumer feels responsible for the share of social surplus corresponding to his shareholding in order to avoid a situation where the social surplus term overwhelms the profit term for a small shareholder.”). In mathematical terms, this is equivalent to changing the utility function in BHZ by multiplying the social preferences term as well as the investment returns term by 1/𝑟. Under that assumption, investors would vote on operational matters by trading off the effects of the decision on firm value and on social considerations, with the weight on social considerations depending on the strength of their social preferences (which would reflect their financial position in the firm), in much the same way as we characterized aggregate shareholder welfare in Section IV.A above. Which of these models best captures how investors would actually think about binding shareholder proposals on operational matters cannot be derived through purely deductive reasoning but rather is ultimately an empirical question, which we return to below.

A second key assumption of BHZ concerns the effect of diversification on investors’ incentives to become informed about votes. An individual investor’s probability of casting the pivotal vote that determines the outcome goes to zero as they become perfectly diversified, for the same basic reason that their interest in the returns on any particular company’s stock goes to zero. This latter effect of diversification plays a key role in BHZ’s analysis, as we have discussed, but with regard to the former, BHZ assumes that “shareholders will vote as if they were pivotal since this is the only case where their vote matters; in other words, they vote the outcome they would like to occur.”199Broccardo et al., supra note 78, at 3114. But a more consistent view about the effects of portfolio diversification is that there would be no reason for an individual investor to give a moment’s thought or attention to how to vote shares or to potential investment funds’ voting policies, because in the limit an individual shareholder has no effect on the world.200Cf. Brav et al., supra note 161, at 505 (finding a positive empirical relation between a retail investor’s ownership position in a company and the likelihood that the investor casts a ballot at the company’s annual shareholder meeting). The prediction of the model would then not be that each investor acts like a social planner but rather widespread rational investor apathy about shareholder votes and about how funds vote, even for socially minded investors.201And these objections do not exhaust the set of critical assumptions that BHZ relies on for their result. For example, their result also hinges on specific choices about the cost structure of the corporate action being voted on (“adopting a technology”). BHZ assumes that the action entails only fixed costs and has no effect on marginal costs. But if it were to increase firms’ marginal costs, then under perfect competition the result would only obtain if all firms adopted it at once. If some firms do not adopt, then the remaining dirty firms would win the entire market. In turn, in equilibrium consequentialist shareholders would no longer view adopting the technology as actually reducing the externality. Their additional assumption of fixed capacity constraints might avoid this problem to some extent, but that represents still another example of how, in our view, BHZ relies on very strong assumptions.

Perhaps the best evidence for evaluating the predictions of BHZ is from shareholder voting on a major class of operational decisions on which shareholders currently are given a binding vote: mergers. Corporate mergers implicate both investors’ investment returns as well as a range of social concerns, including those stemming from increased market power and with respect to the effect of the merger on various classes of firm stakeholders, such as employees and creditors. The model of BHZ predicts that investor voting on mergers would be based not on their own investment returns but rather on such social issues. In short, shareholders would vote for mergers only to the extent they improved social welfare and against mergers that impaired social welfare, regardless of the financial return shareholders could expect from the merger. It is, of course, a claim that calls into question the need for any oversight of mergers on public policy grounds (for example, through antitrust review) as this work would be accomplished through the shareholder vote.

Not surprisingly, this prediction is belied by the evidence: the main concern among shareholders in controversial merger votes is, to our knowledge, never about market power or the effects on other corporate constituencies but rather about the deal price. As an example, consider Michael Dell’s 2013 leveraged buyout of Dell, Inc. When originally proposed, the deal—like many management buyouts—attracted substantial shareholder opposition based on the concern that shareholders were being offered too low of a price, leading Michael Dell to sweeten the deal by offering a special dividend to shareholders.202See David Benoit & Sharon Terlep, Dell Reaches New Deal with Founder, Wall St. J. (Aug. 2, 2013, 7:34 PM ET), https://www.wsj.com/articles/SB100014241278873246359045786434912332027
54 [https://perma.cc/BJ8F-8S5T].
Deal price is a purely distributive concern that implicates investors’ returns; if shareholders cared only about the social welfare effect of a merger, this distributive concern would be irrelevant. It is difficult to reconcile the centrality of concerns about deal price in shareholder voting about mergers—as opposed to concerns about market power or treatment of other corporate constituencies—and BHZ’s model of voting on operational decisions. In contrast, this outcome is consistent with our analysis in Section IV.A above that the overwhelming driver of shareholder welfare under the SSP view is firm value, not shareholders’ social preferences.

4.  The Benefits of Devolving Control to Stockholders

What then would be the benefits, in terms of improved corporate conduct, of devolving control to stockholders? In our view they would be negligible, for the same basic reasons we gave in evaluating the objective functions under SSP and PVM and their feasibility for corporate managers in Sections IV.A and IV.B. We will not recapitulate all of those arguments here, but in short, the predominant consideration that would drive shareholder voting on operational matters would be firm value, not broader social concerns or portfolio externalities. To the extent shareholders’ social preferences did factor into their voting on operational matters, they would entail a form of stated preferences based on the limited information available to shareholders about the full consequences of the vote on a firm’s operations. As such, they would not serve as a reliable guide to shareholders’ revealed preferences about social issues or to social welfare.

While it might be hoped that such a devolution would at least facilitate low-hanging-fruit improvements to corporate behavior—changes that would attract widespread agreement in society—such issues are those that are most likely to be addressed already by law and public policy. Putting operational matters to a shareholder vote involves deploying a type of political mechanism—what Roberto Tallarita refers to as “stockholder politics”—as an alternative to traditional politics.203Tallarita, supra note 190, at 1701. But by our lights, stockholder politics is likely to be much less protective of broader social interests than traditional politics since corporate stockholders are a subset of the broader polity and this subset of voters owns the claims to the corporate profits that would have to be sacrificed in service of those broader interests.

5.  The Costs of Devolving Control to Stockholders

While the social benefits from devolving control to stockholders would be negligible, the social costs would likely be significant. Those costs would come in three main forms. First, allowing shareholders to propose binding resolutions on corporate conduct would result in substantial distraction of management, which would inevitably be drawn into defending corporate policies against social activists pushing for reforms. As we have emphasized previously, managers have a finite amount of time and attention so that this distraction would result in worse corporate performance over time. Second, devolving control to shareholders risks changes to corporate policy that are likely to reduce the well-being of shareholders and the broader society. That is, one cannot be confident that all successful shareholder interventions would ultimately be in shareholder interests, given the many layers of intermediation between beneficial owners and the shares as well as the limited amount of information shareholders would inevitably have about the full costs and benefits of a proposed change in a firm’s operations in this decision-making environment.204Zohar Goshen and Richard Squire term the costs that occur when investors exercise control “principal costs,” a play on “agency costs.” Zohar Goshen & Richard Squire, Principal Costs: A New Theory for Corporate Law and Governance, 117 Colum. L. Rev. 767, 771 (2017). Similarly, Iman Anabtawi argues that giving shareholders more power over operational matters would distort corporate decisions due to the influence of large shareholders with interests that conflict with shareholders’ interests as a class. Iman Anabtawi, Some Skepticism About Increasing Shareholder Power, 53 UCLA L. Rev. 561, 561 (2006). Finally, as other scholars have noted, turning to shareholder voting in hopes of regulating the production of technological externalities comes with troubling political implications. These include the possibility of chilling the perceived need for systematic legislation and regulation,205See Bebchuk & Tallarita, supra note 2, at 168–73. the effective weighting of shareholders’ policy preferences by their wealth,206Marcel Kahan & Edward B. Rock, Corporate Governance Welfarism, 15 J. Legal Analysis 108, 123 (2023). and the vesting of de facto regulatory power in the hands of a few unelected asset managers given the prevailing distribution of voting power in corporate elections.207See Condon, supra note 85, at 8 (“Beyond a mere tallying of positive and negative economic outcomes, the role of investor as private regulator should raise concerns about the compatibility of concentrated corporate control with democratic society—concerns dating back at least as far back as Adolf Berle and Gardiner Means.”); Dorothy S. Lund, Asset Managers as Regulator, 171 U. Pa. L. Rev. 77, 77–78 (2023) (arguing that asset managers effectively supply regulation on matters pertaining to social and environmental matters and highlighting the lack of democratic accountability and government oversight for their policymaking).

V.  THE FUTURE OF CSR IS ESV

Shareholder governance holds significant promise for improving corporate social responsibility. But this promise does not stem from any innovation in our basic understanding of shareholders’ interests along the lines of shareholder welfarism. Indeed, we have argued that changing the corporate objective in the ways urged by shareholder welfarism would fail to meaningfully improve corporate conduct and might even do the opposite. Rather, the ongoing promise of shareholder governance for CSR stems from the prospect of further reductions in certain agency costs and information problems based on the traditional corporate objective, long-term shareholder value. We suspect that there remain opportunities for corporate management to reform firm policies in ways that both increase shareholder value and improve the firm’s social performance, perhaps by addressing the information and incentive problems of ESV we have discussed. But ESV is often misunderstood in the law-and-economics literature. In this final part we begin by addressing those misconceptions and clarifying what we believe to be the most useful understanding of ESV. We then briefly describe an episode at ExxonMobil that illustrates recent innovations in the use of ESV arguments by market actors and the potential promise that ESV holds for advocates of CSR. We conclude this part by identifying a set of key questions about ESV that we think form an important research agenda for the field.

A.  Clarifying ESV as a Concept

Despite its surging popularity in the business world, ESV has received little sustained analysis in legal scholarship. What attention it has received from legal scholars largely reflects one or both of two misconceptions about ESV that we seek to clarify here.

First, some shareholder primacy theorists misconceive ESV as an alternative to traditional shareholder value as a corporate objective.208See, e.g., Bebchuk et al., supra note 42, at 732; Lund, supra note 9, at 94 (contrasting the “traditional” shareholder wealth maximization standard with the “enlightened shareholder value standard”). Relatedly, some CSR-oriented scholars treat ESV as a form of stakeholderism that ultimately requires corporate actions that sacrifice shareholder wealth to further stakeholder interests. Virginia Harper Ho, “Enlightened Shareholder Value”: Corporate Governance Beyond the Shareholder-Stakeholder Divide, 36 J. Corp. L. 59, 98 (2010) (“[I]t is in the cases . . . where market forces pressure firms away from social responsibility–that the contrast between shareholder wealth maximization and enlightened shareholder value is clearest. These are cases where a course of action that maximizes profits imposes negative externalities on stakeholders . . . . If permitted by law, such decisions are fully compatible with a shareholder wealth maximization approach. Under an ESV decision rule, in contrast, the firm must assess the potential impact on stakeholders. If a course of action is optimal only when the costs to stakeholders are ignored, then it should not be taken or the firm must absorb the costs.”). This is not what we refer to as ESV in this Article. For example, in a recent paper Lucian Bebchuk, Kobi Kastiel, and Roberto Tallarita examine “the view that corporations should replace their traditional purpose of shareholder value maximization (SV) with a standard commonly referred to as ‘enlightened shareholder value’ (ESV).”209Bebchuk et al., supra note 42, at 732. After arguing that SV and ESV are operationally equivalent, they conclude that “replacing SV with ESV should not be expected to produce benefits for either shareholders or society.”210Id. at 3.

But their framing of ESV as an alternative corporate objective is, in our view, a category mistake. ESV is not an alternative corporate objective. The enlightenment that ESV calls for involves not an adjustment of the corporate objective itself but rather in how to seek it. ESV is best understood as a reform agenda targeting a particular class of agency costs and information problems that harm not only shareholders but also other corporate stakeholders. Just as one might usefully analyze problems with the design of executive compensation as a distinctive manifestation of and contributor to managerial agency costs,211See, e.g., Lucian Bebchuk & Jesse Fried, Pay Without Performance 4-5 (2004). ESV theory identifies a particular class of agency and information problems worthy of study that might point to their own set of interventions.

Why have law-and-economics scholars instead viewed ESV as advancing an alternative corporate objective? This framing of ESV might stem in part from the grammatical structure of the label: “enlightened” is an adjective, modifying “shareholder value.” Another reason—suggested by Bebchuk and coauthors212Bebchuk et al., supra note 42, at 736.—is that some jurisdictions have added explicit language to corporate statutes highlighting the importance of operating in a socially responsible manner to the achievement of shareholder value. For example, the United Kingdom Companies Act provides

A director of a company must act in the way he considers, in good faith, would be most likely to promote the success of the company for the benefit of its [shareholders] as a whole, and in doing so have regard (amongst other matters) to— . . . 

(b) the interests of the company’s employees,

(c) the need to foster the company’s business relationships with suppliers, customers and others,

(d) the impact of the company’s operations on the community and the environment,

(e) the desirability of the company maintaining a reputation for high standards of business conduct.213Companies Act 2006, c. 46, § 172(1) (UK).

But such a provision does not change the corporate objective from maximizing shareholder value. Rather, we suspect that the existence of stakeholderism as a competing conception of corporate purpose may explain the perceived need to add explicit language endorsing such CSR considerations in pursuing long-term shareholder value. After all, many people believe in stakeholderism, which is indeed a fundamentally different understanding of ends, and not just means, of the corporate form. This leads to several phenomena that might in turn justify explicit acknowledgement of ESV considerations in corporate law.

First, when good faith managers sacrifice short-term profits to act more responsibly in ways that further shareholder value, they might be accused of being stakeholderists! Explicit legal endorsement of ESV can reassure all involved that engaging in CSR is often required to further shareholder value. Second, one could interpret explicit ESV legal language as limiting rather than permissive; it can make clear to corporate managers that they should pursue CSR only to the extent that it furthers shareholder value. This is what the Delaware Supreme Court did in the Revlon case (“A board may have regard for various constituencies in discharging its responsibilities, provided there are rationally related benefits accruing to the stockholders.”).214Revlon, Inc. v. MacAndrews & Forbes Holdings, Inc., 506 A.2d 173, 182 (Del. 1986). Finally, stakeholderists often propagate a caricature of shareholder value theory in which fat-cat capitalists squeeze every last penny out of workers and customers, pollute the environment at will, and otherwise act in outrageous ways all in pursuit of immediate profit.215See, e.g., Lynn Stout, The Shareholder Value Myth: How Putting Shareholders First Harms Investors, Corporations, and the Public vi, 3, 7, 11 (2012) (“Conventional shareholder value thinking . . . . causes companies to indulge in reckless, sociopathic, and socially irresponsible behavior . . . . In the quest to ‘unlock shareholder value’ [directors and executives] sell key assets, fire loyal employees, and ruthlessly squeeze the workforce that remains.”). Legal endorsement of ESV helps combat that distorted view of shareholder primacy.

A second misconception about ESV is that it is useless because the behavior of all the key actors in the corporate system is determined by their incentives and so ESV ideas cannot improve it. One version of this critique focuses on the significant extent to which existing corporate governance institutions already provide substantial incentives for management to maximize shareholder value, including through practices that also further stakeholder interests, which raises the question of whether there remain any such opportunities not yet exploited. As Elhauge puts it, “Agitating for corporations to engage in responsible conduct that increases their profits is a lot like saying there are twenty-dollar bills lying on the sidewalk.”216Elhauge, supra note 41, at 744–45.

Quite the contrary. For one, the mechanisms posited by ESV often involve substantial uncertainty as to how best to maximize long-term shareholder value.217Edmans, supra note 43, at 60. That uncertainty is in part a function of the long time horizon over which the firm will receive the ultimate financial benefits of socially responsible conduct. In contrast, the financial costs of such practices are typically both immediate and certain. As a result, there is no reason to think that all such positive NPV investments in social responsibility will be exploited. In many cases, firm managers will simply make mistakes in striking these uncertain intertemporal trade-offs. These mistakes, moreover, might be systematically biased toward social irresponsibility, given the asymmetry that poses certain, immediate costs against uncertain, future benefits of more responsible conduct.218To be clear, the existence of such a systematic bias is not self-evident, nor is it fundamental to our argument. All that is necessary to make ESV of interest is that there exist unrealized opportunities to reform corporate policy in ways that further both shareholder interests and CSR, not that there are more such cases than there are cases in which corporations engage in excessive CSR from a shareholder value perspective. More fundamentally, management might face conflicts of interest that produce agency costs in the form of inefficiently irresponsible corporate conduct.219Note that this can be the case even when there are other conflicts of interest that might result in management sometimes acting excessively responsibly from a shareholder value perspective. ESV as we define it focuses on eliminating inefficient corporate irresponsibility. One could imagine another reform agenda that focuses on eliminating inefficient corporate responsibility, which we might term “anti-stakeholderism.” In principle these two reform agendas need not be in conflict with one another. As we have explained, the ESV approach is best understood as largely involving concern about a genus of agency costs in the short-termism family.220See infra Section IV.B.1. The key conceptual challenge for ESV theory is thus not how to explain all the cash on the sidewalk but rather to identify governance reforms or other interventions that might realistically reduce these agency costs and produce more cash.

In that vein, a second version of this critique of ESV takes a glass-half-empty perspective on management incentives. For example, Bebchuk and his coauthors argue that, to the extent that managers fail to engage in shareholder-value-maximizing CSR due to incentive problems that lead to short-termism, ESV offers no way out. As they put it: “[A]s long as corporate leaders have short-term incentives, pontificating to them about the importance of taking into account long-term effects, either in general or with respect to stakeholders in particular, would not address short-termism problems.”221Bebchuk et al., supra note 42, at 748.

Their claim exemplifies what economists have termed the “determinacy paradox.”222Brendan O’Flaherty & Jagdish Bhagwati, Will Free Trade with Political Science Put Normative Economists Out of Work?, 9 Econ. & Pol. 207, 208 (1997). This problem arises when an analyst has a positive model of the actors in a system that generates predictions about how those actors will behave, but then nonetheless engages in normative arguments about how those actors should behave.223Id. at 208. If the analyst believes that the actors’ behavior is pinned down by the positive model, what exactly is the point of the normative arguments? That is the logical structure of Bebchuk and his coauthors’ critique, and it does indeed pose an important challenge for ESV theory.

But note that, as a preliminary matter, this basic challenge for ESV theory is shared by all normative arguments in corporate law scholarship. Economic analysis of corporate law relies on a rich set of positive models that explain the behavior of key actors in the system—officers, directors, shareholders, and the like. But in addition to all of their positive theorizing, corporate law scholars have a decidedly reformist bent. After diagnosing some set of pathologies in the corporate system, generally with the aid of a positive model, the typical scholarly article about corporate law then turns to reform proposals that aim to remedy the problem.224See, e.g., Lucian A. Bebchuk, The Case for Facilitating Competing Tender Offers, 95 Harv. L. Rev. 1028, 1030 (1982) (“[F]acilitating competing tender offers is desirable both to targets’ shareholders and to society.”); Lucian Arye Bebchuk, The Case for Increasing Shareholder Power, 118 Harv. L. Rev. 833, 837–38 (2005) (“Part III presents the case for giving shareholders the power not only to elect and replace directors, but also to initiate and adopt rules-of-the-game decisions to amend the corporate charter or to reincorporate in another jurisdiction . . . . [It] also provides empirical evidence of management’s ability to avoid rules-of-the-game changes that are viewed as value-enhancing by a majority of shareholders.”). But if all of the relevant decisionmakers’ behavior is pinned down by incentives, what is the point of this pontificating? If the positive model is right, then why would managers or directors, for example, care about the analyst’s normative arguments? This is a challenge even for normative arguments about what the law should be, since positive models in corporate law scholarship purport to explain even the content of corporate law itself, for example as the inevitable outcome of state competition for charters.225See, e.g., Roberta Romano, The State Competition Debate in Corporate Law, 8 Cardozo L. Rev. 709, 712–25 (1987) (reviewing positive models of state corporate law based on competition for corporate charters). The generality of this analytic challenge for normative arguments in corporate law scholarship has not previously been recognized.226In contrast this challenge has been discussed extensively in public law scholarship. See, e.g., Eric A. Posner & Adrian Vermeule, Inside or Outside the System?, 80 U. Chi. L. Rev. 1743, 1749 (2013) (arguing that public law scholarship commonly suffers from the determinacy paradox insofar that it combines “pessimism about diagnoses with unexplained optimism about solutions”).

Are all normative arguments about corporate governance hopeless then? Thankfully, no. The way out of the paradox is to identify some set of actors that might ultimately be persuaded by the normative argument. The ability to persuade an actor in turn typically requires that the actor have both something to learn and incentives that align to some degree with the recommendation.227O’Flaherty & Bhagwati, supra note 222, at 215. Rather than leading to normative nihilism, the determinacy paradox should instead discipline us as corporate law scholars to be more explicit about the audiences we have in mind for our normative arguments and to explain why—despite our rich positive models—those arguments command attention. We need an unmoved mover in the system who might be open to the normative argument in order for it to make a practical difference.

Two key audiences who often play that role in corporate law scholarship, more or less explicitly, are institutional investors and government officials. To give one illustrative example, consider Lucian Bebchuk and Jesse Fried’s incisive book on executive pay.228Bebchuk & Fried, supra note 211. They argue that a range of common practices in executive pay stem from, and contribute to, managerial agency costs.229Id. at 45–95. For this analysis to deliver a practically useful normative payoff, however, requires there to be an audience for their arguments that might be influenced in such a way that the design of executive compensation improves. The authors argue in part that “[t]his is an area in which the very recognition of problems may help alleviate them,” asserting that “[m]anagers’ ability to influence pay structures depends on the extent to which the resulting distortions are not too apparent to market participants—especially institutional investors.”230Id. at 12. But they also advocate policy changes that would shift power from boards to shareholders, arguing that

[f]or there to be changes in the allocation of power between management and shareholders, investors’ demand for them must be sufficient to outweigh management’s considerable ability to block reforms that chip away at its power and private benefits. This can happen only if investors and policymakers recognize the substantial costs that current arrangements impose—as well as the extent to which solving existing problems requires addressing the basic problem of board unaccountability. We hope that this book will contribute to such recognition.231Id. at 216. But at times the authors leave the identity of the policymaker being appealed to unspecified. See id. at 213. For example, after pointing out that “states seeking to attract incorporating and reincorporating firms have had incentives to give substantial weight to management preferences, even at the expense of shareholder interests,” the authors write,

Giving shareholders the power to initiate and approve by vote a proposal to reincorporate or to adopt a charter amendment could produce, in one bold stroke, a substantial improvement in the quality of corporate governance. Shareholder power to change governance arrangements would reduce the need for intervention from outside the firm by regulators, exchanges, or legislators.

Id. But the identity of the policymaker who they hope will do the “giving” is left unspecified. See id.

The determinacy paradox strikes us as easier to surmount for normative arguments in ESV theory than it typically is in corporate governance theory more generally. After all, ESV theory, by definition, pushes for reforms that are in the interests of both shareholders and other stakeholders so that multiple classes of actors in the system have interests that are to some degree aligned with the reform to corporate practice being urged and might therefore play a role in helping to bring it about.

Normative ESV arguments by academics, for example, might usefully target a range of audiences in the corporate system. Consider Alex Edmans’s 2020 book, Grow the Pie, which seems primarily aimed at teaching managers how focusing on the social value created by the firm is a surer path to shareholder value creation than seeking shareholder value directly.232Edmans, supra note 43, at 23–37. The book provides a lucid account of the relevant empirical literature on these issues that we suspect has important lessons for managers and independent directors. Institutional investors might also benefit from his analysis and be persuaded to adjust their approach to using ESG factors in their investment process. This could well be an area in which clearer recognition of the agency cost problems that deter managers from considering social value may help alleviate them, as Bebchuk and Fried assert about executive compensation.233Bebchuk & Fried, supra note 211, at 12. And to the extent that failures to exploit all opportunities to engage in CSR in ways that benefit stockholders stem from mistakes due to limited information, the potential for ESV arguments to make a difference is even more straightforward.

In sum, the Panglossian argument that nobody could possibly have a useful new idea along the lines of ESV because if it were incentive compatible to adopt a practice that improved CSR in ways that benefit shareholders, corporations would already be doing it, proves too much. As well, as a positive matter, the increase in the use by various actors in the corporate system of normative arguments about corporate practices that sound in ESV terms is by our lights a phenomenon worth studying rather than simply dismissing. Consider, for example, the ESV argument advanced by Blackrock’s Larry Fink in his 2022 Letter to CEOs: “In today’s globally interconnected world, a company must create value for and be valued by its full range of stakeholders in order to deliver long-term value for its shareholders.”234Fink, supra note 101. The audiences for this argument include independent directors, managers, and other investors.

More concretely, the 2021 activist intervention at ExxonMobil by the hedge fund Engine No. 1 similarly illustrates the potential promise ESV holds for CSR. In the spring of 2021, Engine No. 1 initiated a proxy fight based on a platform that was heavily critical of the Exxon’s failure to grapple with the reality of a rapidly decarbonizing world.235For the history of Engine No. 1’s proxy fight, see Jessica Camille Aguirre, The Little Hedge Fund Taking Down Big Oil, N.Y. Times Mag. (June 23, 2021), https://www.nytimes.com/2021/
06/23/magazine/exxon-mobil-engine-no-1-board.html [https://perma.cc/5N2J-CBFD]. From the start, Engine No. 1 emphasized the central importance of climate change and decarbonization for the campaign. As it stated in its opening salvo to Exxon, “It is clear . . . that the industry and the world it operates in are changing and that ExxonMobil must change as well.” Engine No. 1 LLC, Letter to the Board of Directors, Reenergize Exxon (Dec. 7, 2020), https://reenergizexom.com/materials/letter-to-the-board-of-directors [https://perma.cc/6R2G-32HP].
Critically, however, its central argument was that management’s failure to cut back on investment in oil production was bad for business, not just bad for the earth.236Exxon Mobil Corp., supra note 235. As the fund emphasized when it launched its campaign, the company’s total shareholder return over the past ten years had been -20%, compared to 277% for the S&P 500, and it also trailed its industry peers. Id. In its investor presentation, Engine No. 1 argued that the stock’s lackluster performance reflected a fundamental failure at the company to adjust its business strategy to account for long-term demand uncertainty for oil and gas. In particular, Exxon’s long-term business planning “centered narrowly on projections of oil and gas demand growth for decades,” see Exxon Mobil Corp., Proxy Statement (Schedule 14A) 21 (Mar. 15, 2021), leading it to pursue “aggressive capital expenditure plans to chase production growth” that have left “ExxonMobil far more exposed than peers to demand declines,” id. at 9. Additionally, Engine No.1 emphasized that the company’s “refusal to accept that fossil fuel demand may decline in decades to come has led to a failure to take even initial steps towards evolution.” Id. at 6. In this regard, Engine No. 1 excoriated the company for its “total reliance on [the] hope of carbon capture to preserve [its] business model,” id. at 21, which had caused the firm to lack any “credible plan to protect value in an energy transition,” id. at 14. This failure to grapple with transition risk was in contrast to its peers who “have shown it is possible to begin gradually diversifying – and embracing long-term total emissions reduction targets – while maintaining focus on core business profitability.” Id. at 27. However, with a stake amounting to a mere 0.02% of Exxon’s shares outstanding,237Matt Phillips, Exxon’s Board Defeat Signals the Rise of Social-Good Activists, N.Y. Times (June 9, 2021), https://www.nytimes.com/2021/06/09/business/exxon-mobil-engine-no1-activist.html [https://perma.cc/XM32-J3YY]. Engine No. 1 had to win the votes of other institutional investors in order to succeed. In this regard, it reflected precisely the type of challenge faced by proponents of ESV ideas: namely, how could it convince other investors that Exxon was somehow failing to see how its existing policies were destroying long-term shareholder value? Consistent with our analysis of the limits of ESV, the answer was through highlighting a lack of information238For instance, Engine No. 1 argued that the “[b]oard of ExxonMobil will be addressing the most important questions facing the energy industry for years to come,” Exxon Mobil Corp., supra note 235, at 73, but stunningly, not one of ExxonMobil’s independent directors had any prior energy industry experience, id. at 19 (“Prior to our campaign, ExxonMobil’s Board had no independent directors with [prior] energy experience.”). It was for this reason that Engine No. 1 advanced a director slate that could provide the expertise that it believed the “[b]oard has been missing – directors with diverse yet highly relevant backgrounds who have successfully tackled energy industry challenges and bring decades of experience in conventional and alternative forms of energy to help best position ExxonMobil for greater long-term value creation.” Id. at 73; see also FAQs, Reenergize Exxon https://reenergizexom.com/faqs [https://perma.cc/3SW5-FS9T] (“The four highly qualified, independent individuals we have identified can bring to the ExxonMobil Board much-needed experience in value-creating, transformational change in the energy sector.”). and a lack of incentives239For instance, Engine No. 1 criticized the company’s compensation plans for creating “misaligned incentives.” Exxon Mobile Corp., supra note 235, at 57. It also emphasized the inverse relationship between management compensation and stock performance, arguing that the “[d]isconnect results in part from compensation plans that can reward volumes over sustainable value.” Id. at 59. In contrast to its peers, Engine No. 1 noted that ExxonMobil provided little disclosure regarding how managers were held accountable for cost overruns. Id. Nor did the company follow its peers in utilizing a management scorecard with “well defined weights for metrics and targets” that were tied to energy transition risk. Id. at 60; see also id. at 70 (providing examples of “many peer compensation metrics [that] have evolved to incentivize management to create value by looking at the energy transition as an opportunity”). Instead, the company often resorted to “ad hoc” changes to its compensation plans to encourage investment. Id. at 60. As a result, Engine No. 1 argued, “In the same way that ExxonMobil’s changes to incentive plans to reward production led to a focus on growth even as returns declined, we believe the lack of material energy transition metrics could discourage a focus on the future.” Id. at 70. among Exxon’s management. In the end, its message resonated with a critical audience of institutional investors,240Phillips, supra note 237 (“The tiny firm wouldn’t have had a chance were it not for an unusual twist: the support of some of Exxon’s biggest institutional investors.”). Many of these investors expressly acknowledged the ESV-oriented arguments advanced by Engine No. 1. For instance, in statements explaining their support for the dissident board candidates, institutional investors concurred with Engine No. 1’s critique of the company’s performance, particularly its approach to capital allocation, and its “long-term financial underperformance” relative to its industry peers. Cal. Pub. Emp.’s Ret. Sys., SEC Shareowner Alert – Notice of Exempt Solicitation (Form PX14A6G) 1 (May 10, 2021); State St. Glob. Advisors, 2021 Proxy Contest: Exxon Mobil Corporation (XOM) 1 (2021). Investors also expressed concern about the “board dynamics” highlighted by Engine No. 1, particularly its lack of information, with Vanguard highlighting “concerns about the lack of energy sector expertise in its boardroom,” Vanguard Grp., Inc., Voting Insights: A Proxy Contest and Shareholder Proposals Related to Material Risk Oversight at ExxonMobil 2 (2021), https://corporate.
vanguard.com/content/dam/corp/advocate/investment-stewardship/pdf/perspectives-and-commentary/
Exxon_1663547_052021.pdf [https://perma.cc/JH6Z-5DLT], and BlackRock stating the board would benefit from “the addition of diverse energy experience,” BlackRock, Vote Bulletin: ExxonMobil Corporation 4 (2021), https://www.blackrock.com/corporate/literature/press-release/blk-vote-bulletin-exxon-may-2021.pdf [https://perma.cc/DRT4-VPA5]. The incentives argument was also referenced, though not as explicitly as in Engine No. 1’s critique, with Vanguard alluding to “questions about board independence” that it had raised with Exxon for a number of years. Vanguard Grp., Inc., supra, at 2. Several investors also commented on Exxon’s failure to plan adequately for the energy transition and the long-term value of Exxon. For example, in its statement, BlackRock noted that “Exxon and its Board need to further assess the company’s strategy and board expertise against the possibility that demand for fossil fuels may decline rapidly in the coming decades,” adding that the company’s “current reluctance to do so presents a corporate governance issue that has the potential to undermine the company’s long-term financial sustainability.” BlackRock, supra, at 3. Likewise, Vanguard explained that it grounded its “assessment on how any changes to the board’s composition would affect [Exxon’s] ability to oversee risk and strategy and ultimately lead to outcomes in the best interest of long-term shareholders.” Vanguard Grp., Inc., supra, at 2.
allowing Engine No. 1 to win a contested director election to place three new directors on the board of ExxonMobil.

To be clear, we are not arguing that Engine No. 1 was correct in its critique of Exxon’s management on shareholder value grounds. Exxon’s management heavily disputed that claim, and we remain agnostic. Our claim instead is that the intervention was framed in ESV terms, and the key deciders—large institutional investors—appear to have evaluated Engine No. 1’s candidates based on shareholder value considerations.

B.  A Research Agenda for ESV

We conclude by briefly outlining a set of research questions about ESV that we think would shed light on the ultimate scope for further improvements to CSR through ESV-motivated reforms and that we hope future scholarship will address.

First, how big is the gap between perfect ESV behavior (that is, fully realizing all opportunities to further stakeholder interests that also benefit shareholders) and actual corporate behavior with respect to various social issues? In some areas it may be that calls for reforms to corporate practices, even though ostensibly based on ESV considerations, are actually better understood as stakeholderist in nature. It may be that public policy is a better tool for responding to those cases than appeals for CSR. But in other areas there may be substantial scope for further improvements to corporate practice on ESV grounds.

Second, what are the main reasons that corporations fail to realize ESV opportunities? Investigating past episodes of reform to corporate conduct might reveal the extent to which such failures stem from lack of information versus incentive conflicts. For example, has recent empirical research documenting the firm value generated by treating workers well241See Edmans, Does the Stock Market Fully Vale Intangibles?, supra note 62, at 623; Edmans, The Link Between Job Satisfaction and Firm Value, supra note 62, at 9–11. led to the spread of such practices in the corporate world? Diagnosing the underlying causes of failure to engage in CSR in ways that benefit shareholders might in turn provide insights into how to intervene in the system to improve corporate performance.

Third, and relatedly, what are the contours of the ESV reform agenda with regard to interventions and corporate governance reforms that might improve CSR in ways that further shareholders’ interests? For example, to what extent do governance reforms intended to encourage longer time horizons in management decision-making affect CSR behavior? How can executive compensation arrangements advance ESV considerations? Do popular ESV-oriented interventions—such as enhanced climate disclosures, creating board risk oversight or “sustainability” committees,242Lynn S. Paine, Sustainability in the Boardroom, Harv. Bus. Rev., July–Aug. 2014, at 88; Lisa M. Fairfax, Board Committee Charters and ESG Accountability, 12 Harv. Bus. L. Rev. 371, 386–95 (2022). and appointing independent directors with broader experiences—actually affect CSR decision-making?

Fourth, who exactly are the key actors who might be persuaded by ESV arguments for reform to corporate practices? To what extent are managers, independent directors, and institutional investors persuadable on different ESV issues to act to further such reforms?

CONCLUSION

At the turn of the twenty-first century, leading commentators announced an “end of history for corporate law,” declaring that “[t]here is no longer any serious competitor to the view that corporate law should principally strive to increase long-term shareholder value.”243Henry Hansmann & Reinier Kraakman, The End of History for Corporate Law, 89 Geo. L.J. 439, 439 (2000). Yet the two decades since have witnessed continued developments in corporate law theory and practice that seek to find new pathways for generating more socially responsible corporate behavior. These include new shareholder-centric perspectives that go beyond shareholder value and focus managers instead on more holistic conceptions of shareholder welfare. And even within the traditional paradigm of shareholder wealth maximization, promising innovations abound, including in ways that might improve broader social outcomes. All of these developments suggest to us that the history of corporate law has not yet been fully written, and in this Article, we have tried to assess aspects of this latest chapter. Despite the seeming appeal of conceptualizing shareholder interests in broader terms, on closer examination shareholder welfarism offers little hope for improved corporate conduct. Rather, for those seeking to promote corporate social responsibility, the way forward is through a more thoroughgoing, dare we say enlightened, pursuit of shareholder value.

97 S. Cal. L. Rev. 417

Download

* W.A. Franke Professor of Law and Business, Stanford Law School.

† Robert B. McKay Professor of Law, New York University School of Law. For helpful comments and discussions, we are grateful to Emiliano Catan, John Donohue, Alex Edmans, Jill Fisch, Stavros Gadinis, Jeff Gordon, Oliver Hart, Marcel Kahan, Louis Kaplow, Lewis Kornhauser, Zach Liscow, Dorothy Lund, Veronica Martinez, Curtis Milhaupt, Michael Ohlrogge, Frank Partnoy, Elizabeth Pollman, Ed Rock, Roberta Romano, Holger Spamann, Michael Simkovic, Jeff Strnad, Anne Tucker, Luigi Zingales, Jonathon Zytnick, and seminar participants at Columbia Law School, Cornell Law School, Duke University School of Law, Fordham University School of Law, Georgetown University Law Center, Harvard Law School, Hebrew University, NYU School of Law, Stanford Law School, University of Michigan Law School, USC Gould School of Law, the UC Berkeley/Duke Organizations and Social Impact Conference, and the Corporate Law Academic Workshop Series. Ginger Hervey provided outstanding research assistance.

The Discriminatory Religion Clauses

The Supreme Court’s decision in Carson v. Makin is the third in a trilogy of cases dramatically upending the meaning of the First Amendment’s Religion Clauses. Beginning with Trinity Lutheran v. Comer in 2017 and followed by Espinoza v. Montana Department of Revenue in 2020, the Court has moved forward with an aggressive project of transforming the Religion Clauses into a broad anti-religious-discrimination clause. In this paper, I trace this doctrinal devolution and argue that the Court’s novel reinterpretation is deeply misguided. By design, the Religion Clauses require discrimination—religion is to be treated differently from non-religion in a broad range of state action. The contemporary Supreme Court, however, has inverted this most basic insight. The Court’s new Religion Clause jurisprudence is also on a collision course with its burgeoning government speech doctrine. This doctrine recognizes that in a democratic polity, every policy choice entails paths not chosen. Government must be able to select its own message, and in turn, discriminate against those messages it wishes not to communicate, tempered by accountability at the ballot box. Granted, to say that discrimination is sometimes required under the Religion Clauses and the Government Speech Doctrine is not to say discrimination against religion is always constitutional. Protections against objectionable discrimination remain as vital as ever. The Court’s public forum doctrine, for example, protects free expression of religion from content-based discrimination when the government itself is not speaking. The heart of the Court’s recent Religion Clause decisions, however, is a jurisprudentially backward constitutional mandate that government actively subsidize religious speech to avoid a Religion Clause “discrimination” claim. It is a command that government express ideas it may not wish to express. The Court’s reimagining of the Religion Clauses is inconsistent with the First Amendment’s original meaning, potentially harmful to both government and religion, and in direct tension with the Government Speech Doctrine.

 

What a difference five years makes. In 2017, I feared that the Court was “lead[ing] us . . . to a place where separation of church and state is a constitutional slogan, not a constitutional commitment.” Today, the Court leads us to a place where separation of church and state becomes a constitutional violation.

—Justice Sonia Sotomayor1Carson v. Makin, 142 S. Ct. 1987, 2014 (2022) (Sotomayor, J., dissenting).

INTRODUCTION

The Religion Clauses of the First Amendment require discrimination. Such an assertion may appear counterintuitive in an era prone to viewing subjects of controversy through a lens of equality, but by their very terms the Free Exercise Clause and the Establishment Clause demand that religion be treated differently from other objects of government attention. Today, however, the Supreme Court tells us a different story. Despite the clear language in the Constitution, the Court’s most recent jurisprudence suggests that the religion clauses do something very different than what the words chosen by their framers would suggest.

Beginning in 2017, the Court moved forward with an aggressive project of transforming the Religion Clauses into a broad anti-religious discrimination clause. In this paper, I trace this doctrinal misadventure and argue that the Court’s novel reinterpretation is deeply misguided. This approach, I contend, is precisely backwards. The Religion Clauses are not the Equal Protection Clause. The Court’s conflation of the Religion Clauses with anti-discrimination principles directly contravenes the design and intended function of this critical part of the First Amendment. It is also antithetical to a core principle of popular sovereignty: that a state—and hence, the people—must be able to choose its own priorities and be held accountable for the choices it makes, a key premise underlying the Court’s government speech doctrine.

There are many reasons to find fault in Constitutional doctrine. But whether it is substantive due process and the meaning of the word “liberty” in the Fourteenth Amendment or the right to keep and bear arms in the Second Amendment, such critiques typically boil down to this: the Court is either reading too far into the language of the Constitution or not far enough. It is either finding more meaning then is there, or too little. With the Court’s most recent turn in its religion clause jurisprudence something very different has occurred. Instead of going too far or not far enough, the Court has effectively inverted the very purpose of the Religion Clauses. These clauses, as designed by framers with an understanding of the weighty historical role religion has played in society and governance, carve out religion for a uniquely nuanced, one-of-a-kind treatment. Religion is special. It receives an unusual and distinctive protection from government intervention and is subjected to unusual and distinctive limitations on government support. In between these two constitutional poles established by the Religion Clauses, governments have discretion to make religion-related policy choices. But the unique Janus-faced design of the Religion Clauses sends a clear message: the Constitution requires that religion be treated differently.

Up until 2017, critics of the Court’s religion clause jurisprudence generally fell into the standard camps. They argued, for example, that the Court was restricting too much government activity that “respect[s] an establishment of religion,” as Justice Stewart did in his dissent in Engel v. Vitale addressing a nondenominational school prayer.2Engel v. Vitale, 370 U.S. 421, 444–50 (1962) (Stewart, J., dissenting). Such an exercise, to Stewart, simply did not rise to the level of establishing an “official religion.”3Id. at 450 (Stewart, J., dissenting). Other critics have argued that the Court was not capacious enough in defining what it means to “prohibit” the free exercise of religion, such as Justice Brennan’s dissent in Braunfeld v. Brown, in which he asserted that making a religious practice “economically disadvantageous” should be a sufficient free exercise claim.4Braunfeld v. Brown, 366 U.S. 599, 616 (1961). These cases turned on the unique status of religion––and the extent to which government was treating it differently, as required by the Constitution. And in some contemporary cases, it is still taking this approach, moving the needle much more aggressively than in the past, siding with critics who have supported expansive, and distinctive “free exercise” protection.5See Kennedy v. Bremerton Sch. Dist. 142 S. Ct. 2407 (2022). But as of 2017, the Court also started asking an entirely different, and contradictory question. Inexplicably, differential treatment of religion went from a Constitutional mandate to a Constitutional infraction. 

The first sixteen words of the U.S. Constitution’s First Amendment are straight forward: “Congress shall make no law respecting an establishment of religion, or prohibiting the free exercise thereof . . . .”6U.S. Const. amend. I. The constitutional historian Leonard Levy has asserted that “Nowhere in the making of the Bill of Rights was the original intent and meaning clearer than in the case of religious freedom.”7Leonard W. Levy, The Establishment Clause: Religion and the First Amendment xv (1986). On its face this language prohibits the federal government from making or enforcing laws that do either of two independent things: respect an establishment of religion or prohibit the free exercise of religion. For over three-quarters of a century, this language has been understood to have been incorporated by the Fourteenth Amendment, and thus to apply with equal vigor to the states as to the federal government.8See Cantwell v. Connecticut, 310 U.S. 296 (1940).

In addition to defining precisely what is included in the category of laws “respecting an establishment of religion” or “prohibiting the free exercise,”9Philip B. Kurland, The Origins of the Religion Clauses of the Constitution, 27 Wm. & Mary L. Rev. 839, 856 (1986) (quoting U.S. Const. amend. I, cl. 1). the key interpretive challenge of these two clauses has been their inherent tension. In devising the unique structure of the religion clauses (or, we might say religion clause, singular, to emphasize the interdependence of the anti-establishment and free exercise principles) the framers left behind a distinctive jurisprudential task for courts, incomparable to any other part of the Constitution. What is required or implicitly encouraged by one clause might appear to be prohibited by the other––and in between there may be a zone—what the Court has long referred to as a “play in the joints”10Walz v. Tax Comm’n, 397 U.S. 664, 669 (1970).—where a government may, but is not required, to advance the interests of free exercise or anti-establishment without being prohibited from doing so by the countervailing clause.

The precise contours of the religion clauses continue to be worked out. The drafting history of the religion clauses—particularly, the meaning the framers intended to give to an “establishment of religion”—leaves us with gaping holes in our understanding.11Levy, supra note 7, at 84. One notable area of disagreement in the late twentieth century, for example, has been the debate among jurists and scholars as to whether establishment demands so-called strict separation between church and state or mere nonpreferentialism, that is, not preferring one sect or religion over another.12David Reiss, Jefferson and Madison as Icons in Judicial History: A Study of Religion Clause Jurisprudence, 61 Md. L. Rev. 94, 126 (2002). Various justices on the Supreme Court have long presented differing framings of history in their Religion Clause jurisprudence, confirming that the historical “record does not speak in one voice.”13Id. at 144. But regardless of where one falls in these debates, and however “religion” may be defined, one thing seemingly remained a constant: the religion clauses of the First Amendment single out a thing called “religion” for disparate treatment. While the debate was not definitively settled over precisely where the lines of impermissible establishment or prohibition on free exercise should be drawn, what was clear was that the Constitution established unique lines for religion, prohibiting both governmental favoritism as well as active suppression.

This idiosyncratic Constitutional status of religion vis-à-vis government, which may be seen as a form of mandatory discrimination, is grounded in a set of founding-era philosophical beliefs about the need to protect religion from government and government from religion. As Thomas Jefferson wrote in an 1802 letter to the Danbury Baptist Association, the First Amendment “buil[t] a wall of separation between Church & State.”14Thomas Jefferson’s Letter to the Danbury Baptists (Jan. 1, 1802), https://www.
loc.gov/loc/lcib/9806/danpre.html [https://perma.cc/2D4H-5ZGJ].
And while some scholars have disputed the significance of Jefferson’s famed “wall of separation” metaphor, in 1947 the Supreme Court affirmed Jefferson’s reading in forceful terms. It did not merely agree that “[t]he First Amendment has erected a wall between church and state,” 15Everson v. Bd. of Educ., 330 U.S. 1, 18 (1947). it emphasized that the “wall must be kept high and impregnable.”16Id.

 In 2022 however, the Court issued an opinion that was nothing short of radical. For the Religion Clauses, it was a world turned upside-down. This is not to say that changes had not been on the horizon. Before the recent seismic leap, the Court’s religion jurisprudence had been on a steady retreat from the Jeffersonian vision, particularly as we passed into the new millennium. But Carson v. Makin, capping off a trio of cases that began in 2017, was of a different magnitude.

 Thomas Jefferson’s “wall of separation” between church and state has gone from a route impeded by a barrier “high and impregnable” twenty-five years ago, to one riddled with easily breached fissures shortly thereafter, to an obstruction not merely demolished but replaced and paved over by a wide road—with a shuttle bus travelers are compelled to ride and a fare they are compelled to pay. In Carson the Court did not merely backtrack from its longstanding prohibition on the expenditure of government funds on sectarian schooling; it held, for the first time, that the Free Exercise Clause prohibits a state from not using taxpayer money to fund religious education.17Carson v. Makin, 142 S. Ct. 1987, 2010 (2022). A government, in other words, may be constitutionally obligated to do, what for most of the Court’s jurisprudential history addressing the religion clauses it had been forbidden from doing: paying for religious education.

This was so despite the glaring objection that such funding directly conflicts with a straight-forward, textual reading of the Constitution’s prohibition of any law “respecting an establishment of religion.”18U.S. Const. amend. I. Government may be required to utilize taxpayer funds to pay for religious education in spite of the strong belief of the First Amendment’s framers “that no person, either believer or non-believer, should be taxed to support a religious institution of any kind.”19Everson, 330 U.S. at 12. Government may be compelled to provide an affirmative benefit to religion, despite an absence of evidence that it is in fact “prohibiting” a religion’s free exercise. Indeed, a state may be required to pay for religious education even in the face of its own strong policy reasons for not doing so. How did this happen? And what is the constitutional basis for this revolutionary reformulation?

Carson v. Makin and the two cases leading up to its holding transformed the Free Exercise Clause into an anti-religious-discrimination clause. The Religion Clauses however, were designed to produce the very opposite result, to ensure that religion was treated differently. Religion receives special treatment in the Constitution, precisely because the framers appreciated its unique power. Religion has the ability to inspire, to shape humankind’s deepest and most intimate sense of meaning and well-being, to establish and frame social obligations that supersede or conflict with civic commitments, and to ignite wars, social instability, and bloodshed. In the words of Roger Williams, the theologian and founder of Rhode Island whose religious advocacy for separating church and state left an indelible imprint during America’s colonial era, “[t]he blood of so many hundred thousand souls of Protestants and papists, spilled in the wars of present and former ages for their respective consciences, is not required nor accepted by Jesus Christ the Prince of Peace.”20Mark A. Graber, Foreword: Our Paradoxical Religion Clauses, 69 Md. L. Rev. 8, 9 (2009) (quoting Roger Williams, The Bloudy Tenent of Persecution for Cause of Conscience 1 (Edward Bean Underhill ed., The Society 1848) (1644)).

While a vast sphere of human activity is open to government control, establishing an array of rules determining, for example, the boundaries of criminal and civil conduct, Williams emphasized that religion is different. “God requires not a uniformity of religion to be enacted and enforced in any civil state; which enforced uniformity, sooner or later, is the greatest occasion of civil war, ravishing of conscience . . . .”21Id. at 10. According to historian Leonard Levy, James Madison believed that state “[e]stablishments produced bigotry and persecution, defiled religion, corrupted government, and ended in spiritual and political tyranny.”22Levy, supra note 7, at 55. It was clear that religion, in short, must receive special treatment in its relation to the state.

Carson, however, tells us that this is all wrong. Religion is instead to be treated the same as other human endeavors—at least, for certain purposes. Instead of standing as a mandate for distinctive treatment of religion, the newly reconfigured twenty-first century Religion Clauses, prohibit distinctive treatment. How did the Court justify such a profound—and some might say bizarre—reversal? One explanation is that the Court was drawing on the post-Civil War legacy of the Fourteenth Amendment and modern-America’s strong ethic of opposing inequality and discrimination in its many forms.

Perhaps some justices were also responding to an underlying feeling that religious adherents are looked down upon by societal elites and had not been invited onto the equality train with the same gusto as other identity groups; perhaps this jurisprudential turn was their chance at a ticket. Justice Scalia made such feelings clear in a 2004 dissent when he complained that

[o]ne need not delve too far into modern popular culture to perceive a trendy disdain for deep religious conviction. In an era when the Court is so quick to come to the aid of other disfavored groups, . . . its indifference [to those who dedicate their lives to the ministry], which involves a form of discrimination to which the Constitution actually speaks, is exceptional.23Locke v. Davey, 540 U.S. 712, 733 (2004) (Scalia, J., dissenting).

 An aggrieved Justice Thomas, in his recent Espinoza concurrence, points the finger directly at other justices, lamenting that “this Court has an unfortunate tendency to prefer certain constitutional rights over others . . . The Free Exercise Clause . . . rests on the lowest rung . . . .”24Espinoza v. Mont. Dep’t of Revenue, 140 S. Ct. 2246, 2267 (2020) (Thomas, J., concurring).

It is possible that the Court is taking its cues from grievances such as these. But whatever the motive, the Court had decided, without acknowledging that it was doing so, to completely reimagine the Religion Clauses. In the Carson trio the First Amendment’s religion clauses are framed, not as they have been traditionally construed, as granting religion a unique constitutional status, but as a demand that religion effectively be placed on the same plane as everything else. Granted, this insistence on anti-religious discrimination is not evenly applied. As we shall discuss further, its flattening of the religion clauses is selective. In other contexts, the Court—in the very same term it decided Carson—concluded that a state employee, while acting within his official duties, has special rights of religious expression and practice that he would not possess outside of the religious sphere.25See Kennedy v. Bremerton Sch. Dist. 142 S. Ct. 2407 (2022).

I.  WHY DISCRIMINATION?

Although it may not be commonly acknowledged, government is in the discrimination business. It discriminates every time it “establish[es] Justice, insure[s] domestic Tranquility, provide[s] for the common defence, promote[s] the general Welfare, and secure[s] the Blessings of Liberty.”26U.S. Const. pmbl. All of these ends, eloquently laid out by the founding fathers in the Preamble of the U.S. Constitution, necessarily require America’s government of “we the people” to make choices. There are many routes to realizing, maintaining, and even defining domestic tranquility, the general welfare, and core liberties. And for every policy choice, there are paths not chosen. In a world of scarce resources and fierce disputes over how to allocate those resources to best achieve societal goals, government must not merely decide how much to allocate to particular goals, but which goals are worthy of its energies in the first place.

In a functional and sustainable democracy, this process of discrimination ideally keeps a polity on a trajectory of responsiveness and improvement. Government discrimination allows the state to make discerning choices that take into account an array of complex interests and counter-interests. It allows for action rather than paralysis in light of the needs and pressures coming from a multitude of directions, the often overwhelming and conflicting demands that are part and parcel of having to accommodate a large, diverse, and pluralistic population. Government discrimination in a working democracy means that hard choices will be made; costs will be weighed against benefits. But ultimately, if democracy is functioning in its ideal form, these choices will generally reflect societal values, interests, and goals, while helping correct for the errors of the past as they become evident.

Being “discriminating” thus may be associated with thoughtful, careful, decision-making. And indeed, the Constitution itself not only invites relatively open-ended policy-based discrimination rooted in democratic deliberation, but the document in many places calls for particular kinds of discrimination. It tells us in Article II that we must discriminate against those who are not “natural born” when choosing a president.27U.S. Const. art. II, § 1. While the federal government has the power to “lay and collect Taxes,”28U.S. Const. art. I, § 8, cl. 1. certain kinds of taxes are explicitly verboten (or “discriminated against”) such as tariffs “laid on Articles exported from any State.”29U.S. Const. art. I, § 9, cl. 5. A similar form of discrimination characterizes the Religion Clauses of the First Amendment; religion is explicitly designated as a subject of government regulation to be treated differently from non-religion in a broad range of state action.

Granted, this was not the case under the initial conception devised at the 1787 Constitutional Convention in Philadelphia. The framers’ original notion was one in which the federal government was to be inherently limited to powers enumerated in Article I. Since regulation or establishment of religion was not explicitly included among these powers, a discriminatory carve-out for religion was thought to be superfluous. A bill of rights, including such special treatment for religion, was initially deemed unnecessary because as Madison explained, “[t]here is not a shadow of right in the general government to intermeddle with religion.”303 Jonathan Elliot, The Debates in the Several State Conventions on the Adoption of the Federal Constitution 330 (2d ed. 1836). Other founding era notables however, remained skeptical. Many states conditioned their support of the new charter on a pledge to make the implicit, explicit.31 Steven D. Smith, The Religion Clauses in Constitutional Scholarship, 74 Notre Dame L. Rev. 1033, 1038 (1999). Madison was ultimately persuaded of the merits of this alternative view held by many Anti-Federalists. He became concerned that

under the clause of the constitution, which gave power to congress to make all laws necessary and proper to carry into execution the constitution, and the laws made under it, [congress may be] enabled . . . to make laws of such a nature as might infringe the rights of conscience, or establish a national religion . . . .32Id. at 1039 (quoting James Madison).

Madison realized that even under a regime of limited government in which federal powers are circumscribed by their enumeration in Article I, the government may use its lawful powers in ways yet unanticipated––and that this exertion of power may bleed into religious establishment or the freedom of individual exercise. Because the constitutional structure that limited government power could not be relied upon as the sole guarantor that church and state would be confined to separate spheres, as with other discrete topics, insurance in the form of the Bill of Rights was deemed expedient. And with religion, the remedy was especially distinctive. The two clauses of the First Amendment do not merely single out religion, but do so in an unusual Janus-faced manner suited to the sui generous dilemma that plagued the history of church-state relations. As Justice Robert Jackson explained, “the Constitution sets up [a difference] between religion and almost every other subject matter of legislation, a difference which goes to the very root of religious freedom . . . .”33Everson v. Bd. of Educ., 330 U.S. 1, 26 (1947) (Jackson, J., dissenting).

Religion-related practices receive special discriminatory free exercise benefits exempting them from targeted restrictive governmental regulation that in non-religious spheres would constitute an ordinary part of democratic governance. A wide array of behaviors are targeted by government for distinctive kinds of punishment, prohibition, or penalty, but actions that relate to religion—unlike these other realms of behavior—may not be targeted. They receive a free pass from the Free Exercise Clause of the First Amendment. At the same time that government is generally free to choose to partner with, endorse or incorporate a diverse range of philosophical worldviews, values, or private institutions into its operations, religion may not be among them. Religion is uniquely burdened by the Establishment Clause’s distinctive prohibition on intermingling religion and government.

The two religion clauses simultaneously work together and are at odds with one another. On one hand, they may be said to serve similar ends. “An establishment, Madison argued, ‘violated the free exercise of religion’ and would ‘subvert public liberty.’ ”34Levy, supra note 7, at 168. On the other, they appear in direct tension, one seeming to facilitate religious practice by specially prohibiting government interference and the other seeming to discourage religious practice by specially withholding government largess. The clauses both demand unique benefits for, and impose distinctive burdens on, religion––and sometimes these simultaneous constitutional commands overlap, producing puzzling and uncertain results. It is this apparent paradox that set the stage for a peculiar species of constitutional doctrine, an anomalous and sensitive area of jurisprudence with one common baseline: religion must be treated differently.

II.  DEFINING DISCRIMINATION

The Britannica Dictionary defines “discriminating” as “able to recognize the difference between things that are of good quality and those that are not.”35Discriminating, Britannica Dictionary, https://www.britannica.com/dictionary/

discriminating [https://perma.cc/5UN5-LKX3].
However, discrimination is a word with more than one definition. Discrimination may simply describe the act of “recogniz[ing] a difference between things.”36Discriminate, Britannica Dictionary, https://www.britannica.com/dictionary/discriminate [https://perma.cc/46AX-56WV]. Today, the word “discrimination” is commonly understood as a pejorative. A country founded on egalitarian ideals, with a shamefully inegalitarian past and a present in which identity politics are paramount, has given the word “discrimination” toxic properties. This contemporary understanding aligns with another definition, courtesy of Oxford: to “discriminate” is “to treat one person or group worse/better than another in an unfair way.”37Discriminate, Oxford Learner’s Dictionary, https://www.oxfordlearnersdictionaries.
com/us/definition/english/discriminate [https://perma.cc/PBJ3-LR48].

The First Amendment demands that government treat religion in ways that are arguably both “worse” and “better” than the treatment of other subjects garnering the government’s attention. However, such differential treatment is arguably the epitome of “fairness,” that is if one believes applying clearly stated rules of the U.S. Constitution with principled consistency may generally be understood to be a paradigmatic example of “fairness.” Thus, this latter––pejorative––definition is inapposite to the religion clauses. Yet, merely attaching the word “discrimination” to any government action—whether it be in a political speech, a New York Times op-ed, a Fox News commentary, or a Supreme Court opinion—casts reflexive doubt on that act’s legitimacy. Thus, the irony: use of the phrase “government discrimination against religion”—a constitutional mandate serving the interests of both government and religion—will likely strike the average listener as a nefarious wrong.

There are of course many forms of discrimination that are rightfully prohibited by the Fourteenth Amendment, such as invidious differential treatment based on an individual’s race, gender, or sexual orientation. Other forms of identity-based discrimination, including discrimination rooted in religious animus, may be precluded by statutory anti-discrimination laws. However, the existence of unfair or unjust forms of discrimination—that in some cases are forbidden by the Constitution—should not be used to create the misleading impression that the vital government discrimination required by aspects of the First Amendment is in fact an inherent evil that must be stamped out. Acknowledging that parts of the Constitution require or permit some forms of discrimination does not detract from the continued need (or ability under the law) to combat bigotry.

Granted, in certain contexts, evidence that a government is “discriminating” against or in favor of a particular religion may expose a potential Religion Clause violation. But this is not because the clauses contain a general anti-discrimination principle comparable to the Equal Protection Clause or statutory anti-discrimination law, rather, it is because they demand religion be treated differently from other objects of governmental attention.38See, e.g., Est. of Thornton v. Caldor, 472 U.S. 703 (1985). With other government action, the default is that a democratic state generally must be able to make discriminating distinctions in its policy and enforcement choices.

Discrimination is a baseline for an effective governance. It is the stuff of democratic and legal contestation. Thus, for purposes of this article, I will generally use discrimination in its non-pejorative form—as a mere act of recognizing distinctions between different classes of things resulting in some form of differential treatment. Yes, “discrimination” can be unjust or unfair. “Anti-discrimination” laws and the scrutiny courts apply to invidious discrimination under the Equal Protection Clause of the Fourteenth Amendment have long been directed at such unjust forms of discrimination. However, discrimination can also suggest a kind of discernment that is more typically lauded—such as the ability to distinguish a Matisse or a rigorous scientific study at a top research university from the work produced by a seventh grader in their art or science class. In between the extremes there is enormous room to debate as to whether particular distinctions drawn and differences applied are beneficial or harmful, unfair or justified. And, the Court had historically made such “breathing room” between the discrimination required (or merely allowed) under the Free Exercise Clause and the discrimination required (or merely allowed) under the Establishment Clause, a central component of its religion clause jurisprudence.39See, e.g., Walz v. Tax Comm’n, 397 U.S. 664, 669 (1970).

Nonetheless, words are powerful things. They can be used to manipulate, as well as elucidate. Unfortunately, conflating various definitions of “discrimination,” which is all too common today, may serve the former end. A casual use of the word may create the false impression that particular differential treatment is morally or normatively suspect, when in fact it may be socially desirable––or even a legal requirement. Regretfully, the Supreme Court has gotten in on the act. The Religion Clauses of the First Amendment, very much unlike the Equal Protection Clause, mandate discrimination. Yet, as we shall see, recent religion jurisprudence has mischaracterized “discrimination” as a Constitutional wrong, instead of a Constitutional imperative.

III.  HOW WE GOT HERE: THE LOCKE DISSENT FORESHADOWS A NEW FIRST AMENDMENT

A state may have free reign when it comes to establishing an official state bird, flower, or song, but the Establishment Clause insists that religion is different. That same state may not establish Buddhism or Zoroastrianism or Christianity as its official religion. And the Constitution commands not merely that government shall “make no law respecting an establishment of religion,” it may not prohibit “the free exercise thereof” either.40U.S. Const. amend. I. The state, in the guise of its police powers, may regulate, prohibit, punish, and penalize a full spectrum of human behavior, unless that behavior it is targeting constitutes “an exercise of religion.” While political and constitutional theorists may debate the reasons for this mandatory discrimination––many, including Madison, suggest it serves both the interests of the government and the respective religion that may not be established by the government.41Reiss, supra note 12, at 103. While one might debate the extent and nature of the qualitative benefits Madison foresaw, it is “discrimination” loud and clear.

The Court acknowledged this plain reading of the First Amendment as recently as 2004 in a decision by then Chief Justice Rehnquist. He pointed out, in the context of potential state funding for religious training, that the First Amendment’s unique approach to religion “find[s] no counterpart with respect to other callings or professions. That a State would deal differently with religious education for the ministry than with education for other callings [in other words, that it would discriminate] is a product of these views, not evidence of hostility toward religion.”42Locke v. Davey, 540 U.S. 712, 721 (2004). It was not surprising that his opinion allowing for a selective government scholarship program that excluded theological training read like an exercise in constitutional common sense, with just two dissenters. After all, it had only been two years since the Court, in a controversial 5–4 Establishment Clause decision, first allowed a school voucher program that provided tuition aid to private religious schools to stand.43Zelman v. Simmons-Harris, 536 U.S. 639 (2002).

However, beginning with Trinity Lutheran Church of Columbia v. Comer in 2017, followed by Espinoza v. Montana Department of Revenue in 2020, and most recently, in Carson v. Makin in 2022, the Court radically inverted this natural and widely accepted reading of the Religion Clauses. Admittedly, this novel interpretation of the religion clauses did not appear out of the ether. In that same case in which Chief Justice Rehnquist issued his short ten-page majority opinion rejecting the free exercise inspired demand that the State of Washington pay for a student’s post-secondary religious schooling, Justices Scalia and Thomas dissented and articulated the view that would become the approach of a Court majority beginning in 2017.44Locke, 540 U.S. at 726–34.

For many decades prior to this decision, the Court had interpreted the Establishment Clause as an outright bar on state funding of religious exercise.45Carson v. Makin, 142 S. Ct. 1987, 2012 (2022). However, beginning in the late 1990s, with the case of Agostini v. Felton,46Agostini v. Felton, 521 U.S. 203 (1997).and culminating in Zelman v. Simmons-Harris in 2002, the Court orchestrated what Professor Nelson Tebbe has called a “contemporary turnabout” in its antiestablishment law.47Nelson Tebbe, Excluding Religion, 156 U. Pa. L. Rev. 1263, 1265 (2008). Indirect aid to parents for vouchers to pay for religious education, and even some direct aid to religious institutions, was now a constitutional policy option for legislators across the nation.48Id. at 1266. 

With Zelman, the Court’s religion jurisprudence had just jumped from a world in which government funding of religious education had been presumed to be unconstitutional under the Establishment Clause, to one in which a closely divided Court tenuously held that it was permitted under certain narrow circumstances. In Locke, two dissenters, just two years later, were arguing that such funding was not merely allowed, but required under the Free Exercise Clause, foreshadowing the even more radical changes that were soon to come. Effectively, these two dissenters were arguing—in contravention of the well-established conventional textual reading—that rather than requiring religion be treated differently, the religion clauses instead imposed a broad anti-discrimination mandate. In just a decade and a half, this trial run of the anti-religious-discrimination Free Exercise Clause would transform into the majority view on the Court.

Here was the dissenters’ proposed statement of the rule: “When the State makes a public benefit generally available, that benefit becomes a part of the baseline against which burdens on religion are measured; and when the State withholds that benefit from some individuals solely on the basis of religion, it violates the Free Exercise Clause . . . .”49Locke, 540 U.S. at 726–27 (Scalia J., dissenting). The logic might run as follows: “[P]rohibiting the free exercise” of religion under the First Amendment involves imposing some form of “burden” on such exercise. After all, a “prohibition” imposed by a savvy public official seeking to harm or diminish religion would not typically come in the form of a straight-forward law banning a particular religion or religiosity outright; the more strategically astute approach would be a law that indirectly makes certain elements of a religious practice more difficult, or impossible. Laws with an indirect impact on religion may burden religion. The question then becomes, how do we determine whether there has been such a “burden?” It would seem that to the Locke dissenters, if a “generally available” public benefit is not available to all, we may deem those to whom it is not available, “burdened.”

The dissent utilizes an unexpected dose of post-modern relativistic logic that puts government in the foreground. Their reasoning effectively suggests that it is outside forces—in this case the government—that establish reality for religious practitioners. A burden may be inflicted on religion not just by virtue of what government does to religion, but by virtue of what government does elsewhere. It is as if the dissenters were looking to Article III’s demand that compensation of federal judges “not be diminished during their Continuance in Office,”50U.S. Const. art. III. and reasoning that a change in tax law reducing the mortgage interest deduction is unconstitutional because it makes purchasing a home for a judge more expensive, thereby “diminishing” the relative value of their compensation. The baseline for judging whether free exercise has been burdened is not the unique, longstanding, and deeply rooted practices of the particular religion affected by the government action (or inaction), it is government policy and the relative benefits it provides to various other societal actors. With this peculiar logical maneuver, a constitutional provision that on its face demands religion be treated differently—and protected in ways that other life philosophies or practices are not—is inverted to become one that prohibits religion from being treated differently.

At the same time, the dissent begs the question, what does “generally available” mean? Clearly, all public benefits are subject to rules dictating who is, and who is not eligible. A scholarship fund for post-secondary education will presumably not be available to five-year-olds, nor to those who wish to self-educate in isolation in the woods. The concept of “general availability” requires some sort of limiting principle. If “generally available” simply means that the public benefit at issue is offered in accordance with a relatively fixed non-discretionary rule for some category or categories of non-religious purposes or beneficiaries, this anti-religious-discrimination principle would have virtually limitless application. Considering the ubiquity of government in modern society, it would be an invitation for courts to mandate government-funded religion in virtually all spheres of public life. 

For most of the jurisprudential history of the religion clauses, the Court’s primary challenge, considering the inherent tension between the Establishment and the Free Exercise Clause, has been to craft doctrines determining when, and how much discrimination is required. Must religion be discriminated against when public funds incidentally benefit religious institutions in a way that is comparable to how other (secular) institutions benefit, or only when the funds exclusively target and support a particular religion? Must an anti-discrimination law be discriminatorily applied, exempting hiring and firing decisions by religious organizations from the anti-discrimination mandates that otherwise would apply?51See, e.g., Hosanna-Tabor Evangelical Lutheran Church & Sch. v. Equal Emp. Opportunity Comm’n, 565 U.S. 171 (2012). If so, must such discriminatory exemption apply to just religious ministers, or to all employees of a religious organization? These are the sorts of questions the Court previously asked: to what extent, in what manner, and in what settings do the differential treatment rules of the religion clauses apply to religious organizations and practitioners? The two Locke dissenters inverted the doctrinal question in Religion Clause cases, reframing them as an anti-discrimination mandate.

To critique the dissenter’s approach is not to deny that a violation of the Free Exercise Clause may involve discrimination against a religion or a religious practitioner. A legal ban on Rosary Beads would both arguably prohibit the free exercise of religion for practicing Catholics and at the same time discriminate against Roman Catholicism, treating it differently from other religious practices and secular owners of beaded jewelry. A straight-forward reading of the Free Exercise Clause however, would suggest that it is the prohibition on religious exercise and not the differential treatment that constitutes the constitutional infraction.

 As the Supreme Court has itself emphasized, in the Free Exercise Clause, “[t]he crucial word . . . is ‘prohibit’: ‘For the Free Exercise Clause is written in terms of what the government cannot do to the individual, not in terms of what the individual can exact from the government.’ ”52Lyng v. Nw. Indian Cemetery Protective Ass’n, 485 U.S. 439, 451 (1988). Not only is there no evidence of a general anti-discrimination principle in the text of the Free Exercise Clause, as mentioned earlier, there is an explicit pro-discrimination principle. That is, when a broad legal restriction impacts both religious and non-religious actors, it may be that as to those affected religious individuals “free” religious “exercise” is literally being “prohibited,” entitling them, but not the non-religious affected individuals, to a discriminatory exemption from the law.

Granted, the Court has not been consistent on the question of required accommodations under the Free Exercise Clause. In a 1972 case addressing a state’s compulsory high school education law that was at odds with the practices of a particular religious community, the Court concluded that the Free Exercise Clause demands an exemption.53Wisconsin v. Yoder, 406 U.S. 205 (1972). In contrast, the 1990 case of Employment Division v. Smith suggested that such required differential treatment under the Free Exercise Clause should be construed narrowly.54Emp. Div. v. Smith, 494 U.S. 872 (1990). Then in 2012, a unanimous Court—citing both the Free Exercise and Establishment Clause—concluded that religious institutions are entitled to a ministerial exemption that allows them to fire a teacher of secular and theological subjects, even if such firing would otherwise contravene applicable anti-discrimination law.55Hosanna-Tabor, 565 U.S. at 171. The Court explained that “imposing an unwanted minister . . . infringes the Free Exercise Clause, which protects a religious group’s right to shape its own faith and mission through its appointments.”56Id. at 188.

Regardless of the uneven application over the years, the pro-discrimination implications of the Free Exercise Clause are clear. As O’Connor points out in her Smith concurrence, “A person who is barred from engaging in religiously motivated conduct is barred from freely exercising his religion . . . regardless of whether the law prohibits the conduct only when engaged in for religious reasons, only by members of that religion, or by all persons.”57Smith, 494 U.S. at 893 (O’Connor J., concurring). The free exercise remedy, if it is to apply, would only benefit the religious practitioner—freeing him or her up from an otherwise application restriction—while leaving non-religious individuals burdened. It would, in other words, discriminate between religion and non-religion, treating them differently.

The new anti-religious-discrimination interpretation, in contrast, ignores these basic mechanics of the religion clauses. Scalia’s dissenting opinion in Locke is riddled with surprisingly sloppy reasoning. To support his reading of the Religion Clauses, he draws on an analogy to racial discrimination, yet fails to mention that the Court’s jurisprudence there is rooted in an entirely different part of the Constitution, with completely different language, structure and purpose.58Locke v. Davey, 540 U.S. 712, 728 (2004). The Equal Protection Clause of the Fourteenth Amendment provides that “No state shall . . . deny to any person within its jurisdiction the equal protection of the laws.”59U.S. Const. amend. XIV, § 1. It was not designed with the doctrinally formidable Janus-faced structure (and resulting built-in tension) of the religion clauses—which has led the Court to acknowledge a “play in the joints” between impermissible laws “respecting an establishment of religion” and unconstitutional measures “prohibiting” religion’s “free exercise.”60Walz v. Tax Comm’n, 397 U.S. 664, 669 (1970).

In between the two clauses, in other words, there must be some room for laws that promote anti-establishment values but do not violate free exercise, and vice-versa. This is because laws aimed at avoiding establishment—in the direct sense—will almost invariably diminish free exercise; and laws intended to promote free exercise inevitably move toward establishment. It is a conundrum by design, built upon the Framers understanding of the precarious balance needed to maintain a safe buffer between church and state. The boundaries established by Court doctrine on either side necessitate judicial intervention into matters of religion that are not required of other spheres of government action. Yet, completely disregarding this unique structure of the religion clauses, Scalia instead drew a direct analogy to equal protection. To drive home his point, he argued that “A municipality hiring public contractors may not discriminate against blacks or in favor of them; it cannot discriminate a little bit each way and then plead ‘play in the joints’ when haled into court.”61Locke, 540 U.S. at 728 (Scalia, J., dissenting). But unlike the Equal Protection Clause this is precisely what the religion clauses require––discrimination—a delicate dance between anti-establishment and free exercise in which religion is given special treatment on both ends.

History is riddled with religious wars and instability. The Framers’ innovative formulation in the First Amendment was an attempt to protect the new nation from this same fate. Including only an Establishment Clause would have risked a government so intent on divorcing itself from religion that it would end up stymieing it—generating resentment and potentially violent revolt from passionate religious adherents who felt their free exercise was being choked. Include only a Free Exercise Clause and the danger for government and religion falls on the opposite end of the spectrum; a government openly facilitates and becomes intertwined with religious practice risking its politicization, and the perception (and likely reality) that the state is choosing favorites. Bitterness and backlash among those sects not granted politically favored status would naturally result. As an integrated whole, the two religion clauses were a Goldilocks solution.

As Professor Steven D. Smith observes, “[t]he words . . . ‘establishment of religion’ [and] ‘free exercise’—served to define the substantive area over which Congress was disclaiming jurisdiction.”62Smith, supra note 31, at 1045. It was that simple. There is nothing in the First Amendment demanding that if non-religious governmental benefits are distributed, religious institutions should be entitled to equivalent goodies. Quite the contrary. The Equal Protection Clause of the Fourteenth Amendment and the Religion Clauses of the First Amendment are not the same.

IV.  THE RISE OF “NEUTRALITY”

How then to explain the dissenters’ conflation of principles from these two very different amendments in the Constitution—the Religion Clauses in the First Amendment and the Equal Protection Clause of the Fourteenth? It would seem that Scalia in his Locke dissent was drawing on the “neutrality” principle rooted in certain of the Court’s Establishment Clause decisions. In the seminal 1947 decision Everson v. Board of Education the Court upheld New Jersey’s reimbursement of bus transportation costs to parents sending their children to private schools, including those with a religious affiliation.63Everson v. Bd. of Educ., 330 U.S. 1, 3 (1947). The Everson Court recounted the context in which the Framers’ drafted the religion clauses, stressing that early American settlers sought to escape the compulsion in Europe that they financially support churches favored by the government.64Id. at 8. It emphasized—and included in full in the appendix—James Madison’s Memorial and Remonstrance, a tract written in opposition to a Virginia law that would have imposed a tax on its residents to support the established church.65Id. at 11–12.

Despite ultimately rejecting the Establishment Clause challenge, the Everson Court insisted that “New Jersey cannot consistently with the ‘establishment of religion’ clause of the First Amendment contribute tax-raised funds to the support of an institution which teaches the tenets and faith of any church.”66Id. at 16. It simply found that here, “[t]he State contributes no money to the schools. It does not support them. Its legislation, as applied, does no more than provide a general program to help parents get their children, regardless of their religion, safely and expeditiously to and from accredited schools.”67Id. at 18. Under these circumstances the state was “a neutral in its relations with groups of religious believers and non-believers,”68Id. at 17–18 (emphasis added). not unlike if it were providing police assistance for children crossing the street––some of whom happen to be traveling to or from a religious school.

“Neutrality,” in other words, was a way of distinguishing innocuous general welfare laws that just happen to have, among their many beneficiaries, religious individuals or institutions, from those constitutionally problematic laws that “respect an establishment of religion” by using taxpayer funds for targeted support of religion. If anything, neutrality as used in Everson is about understanding that religion must be treated differently, that while government has broad discretionary power to single-out and benefit all-sorts of respective groups or individuals through the policy distinctions it makes, the one exception is religion. The existence of neutrality (that is, that benefits are provided without regard to the religious status of the beneficiaries) provides support for the conclusion that it is not the kind of law that unconstitutionally respects an establishment of religion. Neutrality suggests that government is not targeting religion qua religion for a specific benefit in violation of the Establishment Clause.

 Neutrality was the principle that Scalia seemed to rely upon when he drew an analogy to equal protection in his Locke dissent, explaining that “[i]f the Religion Clauses demand neutrality, we must enforce them, in hard cases as well as easy ones.”69Locke v. Davey, 540 U.S. 712, 728 (2004) (Scalia, J., dissenting). But, as we have seen, the “neutrality” of Everson is nothing like the general anti-religious discrimination rule the Locke dissenters portray it to be. The fact that the Court has turned to neutrality as a consideration in particular Establishment Clause settings does not transform the Religion Clauses more broadly, and particularly the Free Exercise Clause, into sweeping prohibition of religious discrimination.

The neutrality principle laid out in Everson is one evidentiary standard, among many, for determining whether or not a particular state may be targeting religion in a manner that is inconsistent with the Establishment Clause. Indeed, it is a method of determining when discrimination may be constitutionally required. Considering the fact that most policy choices by government will have some effect on some religious actors, neutrality is simply a device for separating the wheat from the chaff. By providing reimbursement of transportation costs for all schoolchildren—attending secular and religious schools alike—a state is no doubt promoting free exercise of religion. It is making it more affordable for religious parents to freely exercise their religion by educating their children at the religious school of their choice. The question then becomes: under these circumstances does the other religion clause demand discrimination, mandating that religion be treated differently and be denied, unlike the secular schools, this benefit?

Neutrality may be a useful tool in some establishment cases, but it is one that the Court has used only when appropriate, and not with consistency. Indeed, illustrating just how far the Court has moved on religion clause issues, we might observe that Everson itself was a closely contested 5–4 decision. Four dissenters were not convinced that a state should be allowed under the Establishment Clause, as part of a neutral public service program available to all parents, to reimburse families for the cost of sending their children to religiously affiliated schools.

The Locke dissent never explains why, by laying out a standard of “neutrality” in a narrow Establishment Clause context, Everson should now be understood to impose an equality rule under the Free Exercise Clause—requiring the Court to mandate, what in Everson, it just barely allowed. As we shall explore further, while establishment and free exercise may represent two ends of a tension rod, respectively they impose distinct kinds of constraints on government. Scalia, in his Locke dissent, conflates establishment and free exercise.

Justice Gorsuch utilized this conflation to profound effect in his 2022 majority decision in Kennedy v. Bremerton School District.70Kennedy v. Bremerton Sch. Dist. 142 S .Ct. 2407 (2022). There he analyzed a public prayer by a public school coach at a public school event as largely a free exercise issue—whereas in the past the issue would almost certainly have been framed along Establishment Clause lines as an unconstitutional instance of a government official injecting his religion into a school-sanctioned activity. As the smoking gun, Gorsuch points out that “[b]y its own admission, the District sought to restrict [the coach’s] actions at least in part because of their religious character.”71Id. at 2422. It sought to prohibit actions “appearing to a reasonable observer to endorse . . . prayer.”72Id. This was the “gotcha” moment to Justice Gorsuch; a conscientious choice by a school district to comply with the separation of church and state principles articulated in the Establishment Clause becomes damning evidence of a violation of neutrality under the Free Exercise Clause. This is a Religion Clause world turned upside-down.

In Kennedy the Court effectively overruled, indeed inverted, its Establishment Clause precedents recognizing an endorsement test.73Id. at 2427. Public endorsement of religion by government went from prohibited, to prohibited to prohibit. This blowtorch to the Court’s previous jurisprudence, however, cannot alter the fact that the First Amendment, by its very terms, demands discrimination; a state may for legitimate policy purposes designate taxpayer funds to a specific circus school, driver’s education school, agricultural school, or most any other school it deems worthy, except if it is targeting religious education. As we shall discuss in the next Section, consistent with the government speech doctrine, a state has largely unconstrained discretion to choose its own policies and policy messages, except with regard to religion.

V.  THE EMERGING GOVERNMENT SPEECH DOCTRINE

There is some irony in this new muddying of the Religion Clause waters, as the Court has in recent years also moved toward clarification of another part of the First Amendment, one that resonates in the free exercise context: the government speech doctrine. The Free Speech Clause has over time come to incorporate a kind of anti-discrimination principle of its own. Despite reading as a simple across-the-board prohibition that “Congress shall make no law . . . abridging the freedom speech,”74U.S. Const. amend. I. modern free speech case law has come to the realization that the most potent threats to expression come in the form of laws that specifically target (or “discriminate” against) particular content or viewpoints. After all, virtually all laws could be said to impact expression; whether it is blocking traffic on an eight-lane highway, setting private property ablaze, or assaulting a police officer in front of the nation’s capital, if human behavior is observable, it may be framed as expressive. Broad exemptions from criminal and civil accountability merely because the harmful behavior at issue happens to be observable would be intolerable; this was clearly not what the framers of the First Amendment had in mind.

The protection of free expression must have some limiting principle. Thus, the Court has come to differentiate between state attempts to silence particular ideas or ideologies from mere content-neutral “time place or manner” restrictions or regulations directed at harmful behavior that incidentally affects expression. Under the Supreme Court’s free speech doctrine, the former discriminatory treatment of certain content or viewpoints is subjected to a much higher level of judicial scrutiny than the latter—neutral regulations that may in some sense be said to inhibit expression, but without regard to content or viewpoint.75See, e.g., Reed v. Town of Gilbert, 576 U.S. 155, 172–73 (2015). As the end of the twentieth century approached, the Court began to explicitly come to terms with the inverse principle. When it is the government that is doing the speaking, it must have the ability to discriminate.

In a sense, like religion under the Establishment and Free Exercise clauses, “government speech” under the free speech clause is different. It is a democratic imperative that government be able to discriminate in the ideas it conveys. Government must have the ability to choose its own message. It is the culmination of its messages and expressive actions, after all, for which the people hold government to account at the ballot box. Government “speaks” by, among other things, subsidizing particular activities, employing individuals to propagate particular messages, or installing monuments that convey certain ideas.76See Rust v. Sullivan, 500 U.S. 173, 192–93 (1991); Pleasant Grove City v. Summum, 555 U.S. 460, 460 (2009). This is, by its very nature, an exclusionary activity.

As the government chooses to spread one message, it necessarily declines to communicate others. It discriminates based on content or viewpoint. As a new administration takes the helm in response to a shift in voter sentiments, a government will likely change its message. A city government might remove a statue of Robert E. Lee from a public park. It might replace that statue with one depicting the civil rights triumphs of Martin Luther King, Jr. A group of Civil War reenactors may object. However, their recourse is not in a First Amendment that guarantees them a right to have the government send the message they want it to send. It is the political process. The Court made this point succinctly in a 1991 case that would come to be described as the first in a series of cases that form the government speech doctrine.77Helen Norton, The Government’s Speech and the Constitution 32–34 (Alexander Tsesis ed., 2019).

Rust v. Sullivan involved a government program that appropriated public funds for certain family-planning services.78Rust, 500 U.S. at 178. In so doing, Title X of the Family Health Service Act stipulated that none of the allocated funds were to be used in programs that included abortion as a family-planning method.79Id. As a plain-vanilla First Amendment free speech issue, one might assume that the government could not prohibit a counselor or physician from merely discussing a legal abortion as a medical option. Such discriminatory censorship directed at particular content might seem, on the most basic level, antithetical to core First Amendment principles. However, when it is the government that is speaking—as is arguably the case with a government program intended to promote certain goals but not others—the First Amendment prohibition on content or viewpoint-based discrimination is flipped on its head. We expect an anti-abortion administration to “be discriminating” when it comes to the messages it chooses to send about this volatile issue, just as pro-abortion rights voters would expect elected officials who run on a prochoice platform to propagate government speech that facilitates, rather than inhibits, the right to choose. To drive home its point, the Court in Rust provided this example: “When Congress established a National Endowment for Democracy to encourage other countries to adopt democratic principles, . . . it was not constitutionally required to fund a program to encourage competing lines of political philosophy such as communism and fascism.”80Id. at 194.

Thus, if we return to the Court’s anti-discriminatory religion clause innovation, we can see another glaring tension. Even before this current Supreme Court’s most recent Religion Clause turnabout mandating certain government expenditures on religion, some had expressed concern that the growing prominence of the government speech doctrine might diminish previously viable Establishment Clause challenges—because of their potential framing as government speech.81Carol Nackenoff, The Dueling First Amendments: Government as Funder, as Speaker, and the Establishment Clause, 69 Md. L. Rev. 132, 147–48 (2009). But under the Court’s new regime, the anti-religious-discrimination doctrine and the government speech doctrine are on a collision course. A constitutional mandate that government subsidize religious speech to avoid a free exercise “discrimination” claim (just because such subsidy is also available to certain non-religious recipients), is a command that it express ideas it may not want to express, using taxpayer money. It is counter-majoritarian, and directly contradicts the principle underlying the government speech doctrine. In Rust, the Court reiterated the common sense conclusion that “[t]he Government can, without violating the Constitution, selectively fund a program to encourage certain activities it believes to be in the public interest, without at the same time funding an alternative program which seeks to deal with the problem in another way.”82Rust, 500 U.S. at 193. It turns out, however, that this is not the case; that is, at least according to the Court’s novel anti-discriminatory religion clause doctrine.

One might respond, however, as pointed out earlier, that religion is different. Might there be something about religion that would justify a diversion from the otherwise applicable government speech principle? Could it be that this difference merits an exception from the intuitive notion that a government—as a representative of “we the people”—should be able to choose which policies or messages to propagate, and which messages not to endorse, or simply not expend taxpayer resources on? The Constitution, after all, already carves out certain areas in which simple majoritarian politics will not do, requiring instead a super majority for policy change. Fifty-one percent of the population, in other words, cannot do away with probable cause; the Constitution would have to be amended.

One might argue, for instance, that as a fundamental constitutional right, the free exercise of religion should be exempt from the baseline government speech rule. This is quite similar to what was argued by the dissenters in Rust. They pointed to the fact that the right to choose abortion under the “liberty” guarantee in the Fifth Amendment was (at the time) a fundamental constitutional right. As such, selective discrimination against the expression of certain medically pertinent information facilitating that freedom of choice, even under the auspices of a government program, was unconstitutional.83Id. at 216.

 The Court, however, rejected this argument. It also left little room to doubt the basis of this rejection. Citing Regan v. Taxation with Representation, a decision in which the Court upheld a narrowly selective subsidy for lobbying by certain types of organizations, it explained that a “legislature’s decision not to subsidize the exercise of a fundamental right does not infringe the right.”84Id. at 193. Thus, it would seem that the fundamental constitutional rights argument cannot explain the Court’s new anti-religious-discrimination doctrine. The Court’s government speech precedents directly conflict with today’s Court’s characterization of a failure to fund religious education as a “penalty” imposed on that religion.85Espinoza v. Mont. Dep’t of Revenue, 140 S. Ct. 2246, 2255 (2020).

VI.  THE RADICAL TRINITY

If a jurisprudential entrepreneur were on the lookout for an ideal test case to sell a radical reformulation of the Court’s approach to the religion clauses, the facts of Trinity Lutheran Church of Columbia v. Comer would certainly fit the bill. On the surface, this case about the re-surfacing of children’s playgrounds in Missouri involved a highly sympathetic petitioner and addressed relatively un-weighty issues of church and state. To promote recycling and benefit children in low income areas, the state government allocated funds on a competitive basis to help nonprofit daycare centers replace older, harder playground surfaces with ones made from recycled tires.86Trinity Lutheran Church of Columbia, Inc. v. Comer, 582 U.S. 449, 454–55 (2017). Unfortunately for Trinity Lutheran Church, it discovered that its preschool and daycare center were ineligible.87Id. Article I, Section 7 of the Missouri Constitution provided that “no money shall ever be taken from the public treasury, directly or indirectly, in aid of any church, sect or denomination of religion.”88Id at 455. Missouri categorically disqualified religious organizations from receiving grants under the program.89Id.

Although the District Court did not mention the government speech doctrine by name, it upheld the Missouri program using reasoning consistent with the doctrine’s underlying principles. It drew an analogy to a case upholding a state’s “mere” choice not to fund a particular “category of instruction,”90Trinity Lutheran Church of Columbia, Inc. v. Pauley, 976 F.Supp.2d 1137, 1148 (W.D. Mo. 2013). suggesting that it was within Missouri’s discretion to determine the scope of its programs. This includes the choice not to subsidize playgrounds run by religious institutions with public money. In concisely rejecting a free expression argument, the District Court dismissed any notion that the program was designed as an “open forum” for speech.91Id. at 1157.

Consistent with the pro-discrimination implications of the religion clauses, it pointed to the state’s “antiestablishment” interests in preventing religious organizations from receiving government funds.92Id. at 1148. Even if Missouri was not required to promote this interest to the extent it did––prohibiting any receipt of funds by religious organizations––significant “play in the joints” exists between what is prohibited by the Establishment Clause and what is required by Free Exercise.93Id. at 1147. The District Court reasoned that Missouri’s more robust prohibition (what we might certainly call “discrimination” against religion), supports the antiestablishment values built into the religion clauses.94Id. at 1148. Indeed, according to the Court, it would be patently “illogical” to presume that a choice not to fund religion to avoid potential entanglement with government necessarily reflects a hostility toward religion.95Id. The District Court emphasized that the grant here would be paid directly to the religious organization, making the antiestablishment concerns even more compelling than programs designed to sever the direct link between government aid and religious institutions by putting the choice to spend in the hands of private individuals.96Id. at 1152.

The Eighth Circuit affirmed the District Court decision, characterizing the appellant as “seek[ing] an unprecedented ruling—that a state constitution violates the First Amendment . . . if it bars the grant of public funds to a church.”97Trinity Lutheran Church of Columbia, Inc. v. Pauley, 788 F.3d 779, 783 (8th Cir. 2015). In no uncertain terms, it rejected the notion that a state could be compelled to provide taxpayer funds directly to a church: “No Supreme Court case” it explained, “has granted such relief.”98Id. at 784. Moving to an approach in which every generally available public benefit becomes a baseline in which we might scrutinize the denial of comparable benefits to religious actors, would, according to the Circuit Court, constitute “a logical constitutional leap.”99Id. at 785. It would fundamentally recast the Free Exercise Clause from a provision that demands religion be treated differently, to one that prohibits discrimination against it. It would require a repudiation of decades of precedent, and of our foundational understanding of how the religion clauses were to function. In blunt terms, the Circuit Court conceded that “only the Supreme Court can make that leap.”100Id.

But the Supreme Court had indeed changed. Beginning with this unassuming little case about playground surfaces, it was poised to make just such an unprecedented and radical shift in its religion clause jurisprudence. Granted, Chief Justice Roberts, in his majority opinion that overruled the Eighth Circuit in Trinity Lutheran, did not frame his decision in this way. Roberts has developed a reputation for strategic incrementalism, in which the seeds of what will eventually blossom into highly consequential doctrinal change are planted in unassuming soil.101Linda Greenhouse, Justice on the Brink: The Death of Ruth Bader Ginsburg, the Rise of Amy Coney Barrett, and Twelve Months that Transformed the Supreme Court 219 (2021). However, the Trinity Lutheran dissenters did not mince words. Emphasizing the high-stakes of this seemingly low-stakes decision, Justice Sotomayor tells us that “[t]his case is about nothing less than the relationship between religious institutions and the civil government . . . [t]he Court today profoundly changes that relationship.”102Trinity Lutheran Church of Columbia, Inc. v. Comer, 582 U.S. 449, 471–72 (2017) (Sotomayor, J., dissenting).

Roberts’s analysis begins by setting the stage for the Court’s new Free Exercise non-discrimination principle. He cites as a broad rule the rationale of a narrow Free Exercise decision that happened to involve targeted discrimination against a particular religious sect. Granted, the language in the 1993 case Church of the Lukumi Babalu Aye, Inc. v. City of Hialeah103Church of the Lukumi Babalu Aye, Inc. v. City of Hialeah, 508 U.S. 520 (1993).  gave Roberts a good deal to work with. Although the decision was centrally about, as Justice Kennedy explained in the second sentence of the opinion, the “fundamental nonpersecution principle of the First Amendment,”104Id. at 523. it was peppered with the ominous suggestion that impermissible religious discrimination was afoot. However, there is no reason to conclude that the mere relevance of discrimination in this case would convert the religion clauses into a general anti-discrimination rule. Here discrimination simply served as evidence that this particular law should be understood as an unconstitutional prohibition of the free exercise of religion. Like the neutrality principle discussed above, the discriminatory nature of the law was highlighted to demonstrate that things were not as they seemed; a law that may have appeared neutral on its face, was in fact targeting a particular religion’s practices, and thus, quite literally, prohibiting “free exercise” of that religion.

The dilemma with the religion clauses, as with free speech, is that there will necessarily be a vast number of laws aimed at addressing a wide range of social ills that have the subsidiary effect of in-part “prohibiting” the free exercise of particular religions (or “abridging” expressive activity). And the Court has never taken the position, for understandable reasons, that all such laws are unenforceable as to religious practitioners (or to those whose actions are, in part, “expressive”). As Justice Scalia opined, in a country of vast religious diversity, adopting a rule that would strictly scrutinize any neutral, generally applicable law that somehow could be said to intrude on a religious practice would be “courting anarchy.”105Emp. Div. v. Smith, 494 U.S. 872, 888 (1990). The doctrinal parameters of whether, and when, a religious exemption may be required under such circumstances continue to evolve. However, it is clear that laws advancing legitimate, non-religion-related policy ends that incidentally impact the free exercise of certain religious actors are not automatically deemed constitutionally suspect.

No doubt, in drafting the First Amendment the framers sought to prohibit the type of targeted religious persecution that was all too common in the old world.106Babalu, 508 U.S. at 532. But again, the concern was that government not prohibit free exercise through persecution, not that it refrain from treating religion differently from other subjects (something that it is required to do under a straight-forward reading of the text of the First Amendment). Government persecution might be achieved through direct measures that leave little ambiguity as to the intended objective. However, a government intent on punishing, stigmatizing, or driving away an unpopular religious minority might also use non-religion-related policy justifications as a pretext for doing so. It may craft laws that are intended to impede the practices of certain religious believers but justify those laws on legitimate non-religion-related public policy grounds. Or, a legislature might truly have mixed motives. Determining whether or not there has been a free exercise violation under such circumstances may prove difficult. Thus, in this context, identifying “discrimination” may become a vital tool in sussing out whether intentional religious suppression, or a mere side effect of an unrelated policy goal, is occurring.

Preventing animal cruelty was the stated policy goal in Church of Lukumi Babalu. Upon investigation however, this facially legitimate objective was found to have been a front for religious animus. The case involved four ordinances in the south Florida city of Hialeah. Together, they prohibited certain forms of animal sacrifice, a practice associated with the Santeria religion.107Id. at 524–28. The ordinances were apparently spurred on by the imminent prospect of a Santeria church opening in Hialeah and the hostility and discomfort many residents and city council members held toward Santeria and its traditional practices.108Id. at 541–42.

The city argued that its ban on animal sacrifice was justifiable on non-religious grounds. It cited not just protecting animals from cruel treatment, as mentioned above, but also the health risks involved, the emotional injury to children that might result from witnessing such killings, and the interest in restricting slaughter to particular areas of the city.109Id. at 529–30. The narrow ban however, was carefully crafted to exclude virtually all animal killing other than religious sacrifice, and even within this category it exempted kosher slaughter.110Id. at 535–36. The Court concluded that “Santeria alone was the exclusive legislative concern. . . . [K]illings that are no more necessary or humane in almost all other circumstances are unpunished.”111Id. at 536. This was, as Justice Souter pointed out in his concurrence, “a rare example of a law actually aimed at suppressing religious exercise.”112Id. at 564 (Souter, J., concurring).

The Court unanimously struck down the ordinances as a violation of the Free Exercise Clause.113Id. at 546. It was from this unexceptional holding in Church of Lukumi Babalu—prohibiting a legal ban directly targeting practices that were a clear element of the sect’s religious exercise—that the Court in Trinity Lutheran extracts from the Free Exercise Clause a strikingly broad anti- religious-discrimination rule. The new rule requires taxpayer money be used to facilitate the religious mission of an organization—that is, if such funds are available to secular organizations.

Granted, the Church of Lukumi Babalu Court identified, through a close examination of the text of the ordinances at issue and the broader social context, blatantly discriminatory treatment targeting particular practices of a particular religious group. And at times, Kennedy used language to emphasize the significance of such unequal treatment, for example, when he stated that “[a]t a minimum, the protections of the Free Exercise Clause pertain if the law at issue discriminates against some or all religious beliefs or regulates or prohibits conduct because it is undertaken for religious reasons.”114Id. at 532. In this context, this observation simply points out that a restriction on free exercise that is specifically directed toward a particular religion or religious practice is a First Amendment red flag. Such a law presents a sharp contrast to generally applicable laws that affect, and are directed toward, religious and non-religious actors alike. The fact of “discrimination,” in other words, helps courts home in on the most egregious and likely unconstitutional prohibitions on free exercise. Nothing in the decision, however, would suggest that it is the “discrimination” that is the free exercise offense, nor that unconstitutional “discrimination” should be interpreted to encompass a mere choice by a government not to provide financial support to particular religious organizations.

Indeed, Roberts’s reliance upon Church of Lukumi Babalu is particularly curious considering that it was issued just two years after Rust v. Sullivan. As discussed above, this is the seminal government speech case in which the Court explicitly affirmed a government’s power to discriminate—to be selective and make substantive distinctions as to the programs it chooses to fund or not fund.115See supra Part V. What is Chief Justice Roberts’s response to this apparent contradiction? He tells us that “Trinity Lutheran is not claiming any entitlement to a subsidy. It is asserting a right to participate in a government benefit program without having to disavow its religious character.”116Trinity Lutheran Church of Columbia, Inc. v. Comer, 582 U.S. 449, 451 (2017).

But how is claiming a right to receive government largess by “participat[ing] in a government benefit program” that one is not qualified to participate in, anything but an assertion of an “entitlement to a subsidy?”117Id. The Chief Justice’s artful reframing and rephrasing of Trinity Lutheran’s argument does not alter the fundamental facts. After this decision the government in Missouri is required to use taxpayer money to subsidize what on policy grounds it does not wish to subsidize. The Chief’s attempt to sugarcoat its radical decision notwithstanding, the unelected Supreme Court is telling an elected government how it must legislate and allocate its resources—a command that is in direct conflict with its own government speech doctrine.

The only other ostensibly on-point case cited by the Trinity Lutheran Court as support for its innovative religion clause non-discrimination rule was the 1978 plurality opinion in McDaniel v. Paty.118McDaniel v. Paty, 435 U.S. 618 (1978). Under the Tennessee Constitution, clergy were disqualified from serving as state legislators, and thereby not permitted to serve as delegates to a state constitutional convention.119Id. at 620–21. The Supreme Court struck down the exclusion on Free Exercise grounds. The plurality explained that this exclusion of ministers from state legislatures was a practice that was implemented in seven of the original thirteen States. It was instituted “primarily to assure the success of a new political experiment, the separation of church and state.”120Id. at 622.

However, the notion that clergy members should ipso facto be excluded from legislative positions remained controversial. This was so despite the fact that the First Amendment did not at the time apply to the states (it would not be explicitly incorporated until well after the ratification of the Fourteenth Amendment in 1868). Even James Madison, “the greatest advocate for the separation of state and church” 121Andrew L. Seidel, The Founding Myth: Why Christian Nationalism Is Un-American 37 (2019). and primary drafter of the Constitution’s religion clauses suggested (in contrast with Thomas Jefferson’s initial position) that disqualification resembled a kind of unjust punishment reserved for those who happened to choose religious professions.122McDaniel, 435 U.S. at 624. To Madison, the exclusion itself might even constitute a breach of the church-state separation, in that religion was to be exempted “from the cognizance of Civil power.”123Id. at 624. One can thus see the parallel Roberts was attempting to draw with Trinity Lutheran—a law that was arguably “punishing” a playground operator, denying it the opportunity to benefit from a recycled tire resurfacing program, merely due to its religious affiliation.

However, with the help of the government speech doctrine, the distinction between Trinity Lutheran and McDaniel becomes immediately clear. A policy choice as to how the government will use taxpayer dollars—what kinds of interests or schools or playgrounds it will support—is fundamentally different from a law that makes distinctions as to who may legislate in the first place. The former represents a choice as to the policy message the government will communicate, a democratic imperative; the latter represents a choice to exclude certain voices from the possibility of being a part of that government, an anti-democratic exclusion. To suggest that the right to run for office in a democracy is a mere government “benefit” comparable to a government program that helps fund playground resurfacing124Trinity Lutheran Church of Columbia, Inc. v. Comer, 582 U.S. 449, 462 (2017). is to demean a core element of representative democracy. It conflates the ability to select a representative with the naturally selective product of representative democracy; it degrades them both by suggesting that a democracy-affirming Court intervention to prevent limitations on who we may choose as a representative is somehow analogous to a democracy-inhibiting limitation on a government to make policy choices.

McDaniel was also grounded in an individual right to practice one’s religion. The Court explained that “the right to the free exercise of religion unquestionably encompasses the right to preach, proselyte, and perform other similar religious functions, or, in other words, to be a minister of the type McDaniel was found to be.”125McDaniel, 435 U.S. at 626. McDaniel’s right to free exercise was being conditioned upon his surrender of democratic political participation, the choice to run for office. His desire to serve as a delegate to a state constitutional convention was not a request to have the state subsidize his religious activity, except to the extent than any government employee’s private activities might be said to be subsidized by a state salary.

Trinity Lutheran in contrast, involved not an individual’s rights, but the rights of a collective entity. It described its Child Learning Center’s mission as “provid[ing] a safe, clean, and attractive school facility in conjunction with an educational program structured to allow a child to grow spiritually.”126Trinity Lutheran, 582 U.S. at 455. Trinity Lutheran, in other words, was seeking state tax dollars to advance its religious goals as a collective entity. The loss by Trinity Lutheran of the opportunity to participate in a subsidized playground surface program was nothing like the Hobson’s choice that confronted McDaniel. He was not seeking support from the government for his religious works. For McDaniel, under the Tennessee law he was forced to either forfeit his right to fully participate as a citizen or refrain from free religious exercise. 

Chief Justice Roberts finds commonality in McDaniel and Trinity Lutheran, emphasizing the status-based nature of the discrimination in both cases.127Id. at 459. He characterized the policy in Missouri as “expressly discriminat[ing] against otherwise eligible recipients by disqualifying them from a public benefit solely because of their religious character.”128Id. at 462. The McDaniel Court similarly stressed the unique way the law in Tennessee disqualified the petitioner from office “because of his status as a ‘minister’ or ‘priest.’ ”129McDaniel, 435 U.S. at 627. And indeed, the Court has in recent years frequently conflated the individual and the collective; but there can be good reason to acknowledge the differences between the two.

At the heart of classical liberalism is a respect for the individual. The notion that status-based individual deprivations are particularly repugnant is found in many parts of the Constitution itself—whether it is the prohibition on Bills of Attainder,130U.S. Const. art. I, § 9, cl. 3. the demand that “no religious Test shall ever be required as a Qualification to any Office or public Trust under the United States,”131U.S. Const. art. VI. or that the right to vote shall not be denied “on account of race, color, or previous condition of servitude”132U.S. Const. amend. XV. in the Fifteenth Amendment. Although the Court has extended many individual rights in the Constitution to collective entities, there is reason to be skeptical that the same set of concerns applies here.

Tennessee justified its disqualification of a certain category of individuals from elective office on the basis of the “leadership role” and “full time” promotion of “religious objectives” of those who choose to be ministers and priests.133McDaniel, 435 U.S. at 634–35. Citing its goal of maintaining the separation of church and state, the state emphasized its concern that the religious commitments of ministers and priests would at times interfere with their duties as a state legislator.134Id. at 645. Implicit in the plurality decision rejecting this rationale is the understanding that human beings are more than just their chosen avocation. A “unique disability” imposed on an individual because they “exhibit a defined level of intensity of involvement in protected religious activity”135Id. at 632. is, quite simply, highly distinguishable from differential treatment of legal entities based upon their respective, narrowly defined legal purpose.

Nonetheless, the Trinity Lutheran Court finds the organization’s status-based disqualification from the recycled tire playground surface program to be relevant, and sufficiently analogous to the disqualification from office faced by McDaniel. As a result, the Court found Trinity Lutheran merited a similar legal outcome. The Court’s focus on the status-based nature of the religious discrimination at issue also served to distinguish Trinity Lutheran from the 2004 decision Locke v. Davey, the seemingly on-point precedent discussed previously in which the Supreme Court reached the opposite conclusion.

In Locke the Supreme Court upheld a scholarship program in Washington State that, although available for a full range of postsecondary education degrees, stipulated funds could not be used by students “pursuing a degree in devotional theology.”136Locke v. Davey, 540 U.S. 712, 715 (2004). Roberts reasoned that in Locke the student was not denied the benefit of the program on the basis of his religious status, as was true of Trinity Lutheran, but “because of what he proposed to do—use the funds to prepare for the ministry.”137Trinity Lutheran Church of Columbia, Inc. v. Comer, 582 U.S. 449, 464 (2017). Thus, for the Trinity Lutheran Court, the distinction between religious discrimination based on religious “status” and religious “use” appeared to be determinative.

The silver lining of deciding to have the opinion turn on this questionable analogy between the status-based discrimination against the individual minister in McDaniel and the collective religious institution in Trinity Lutheran, is that it established a rule that would, in theory, still allow for government to make crucial policy distinctions consistent with the government speech doctrine. As long as the government is not declining to spend on the basis of religious status, a government might still decline to draw on finite state resources to fund religious action. A government might conclude, for example, that spending on such religious “use” would be unwise, have benefits that are unsupported by evidence, reflect objectives inconsistent with the state’s current policy goals, or simply on balance represent a less weighty spending priority than other competing governmental aims.

This “status” versus “use” distinction, however, would not have staying power. Locke would ultimately be narrowed dramatically, largely relegated to doctrinal irrelevance. In Trinity Lutheran the status/use test was thrown into question in a concurrence by Justices Gorsuch and Thomas. Gorsuch, foreshadowing the Court’s eventual path in Carson, would have distinguished the contradictory outcome in Locke on the basis of its narrow exclusion of scholarship funds for devotional theology and the “long tradition against the use of public funds for training of the clergy.”138Id. at 470 (Gorsuch, J., concurring). For Gorsuch, not only was the status/use distinction likely to be difficult to apply in practice, but it was also irrelevant for the purposes of First Amendment free exercise. The reason? To Gorsuch, “that Clause guarantees the free exercise of religion, not just the right to inward belief.”139Id. at 469.

But this is clearly incorrect. The language of the Free Exercise Clause does suggest a “guarantee.” It no more “guarantees” free exercise than the Free Speech Clause “guarantees” free speech or the Second Amendment “guarantees” that each citizen will be supplied with her own private arsenal. It merely prevents the state from interfering with or “prohibiting,” such freedom. Free exercise of religion may be hampered by friends or family, a wide range of private actors, or the free market itself. Practicing one’s religion may be time consuming, expensive, embarrassing, or stigmatizing. Indeed, it is precisely this kind of interpretive line-blurring of the Free Exercise Clause by Gorsuch that the government speech doctrine rejected when it came to the Free Speech Clause. One is not “guaranteed” an equal opportunity to have the government promote your message of x just because it has chosen to run a public service announcement promoting y.

VII.  ESPINOZA AND THE TRINITY LUTHERAN AFTERMATH

Just three years later in Espinoza v. Montana Department of Revenue the Court broadened the applicability of this fallacious reading. The Montana Constitution included a provision that barred government aid to religious schools.140Espinoza v. Mont. Dep’t of Revenue, 140 S. Ct. 2246, 2251 (2020). Under this “no-aid” provision that the State’s Supreme Court had rejected, a private school tuition assistance program that would have granted “a tax credit to anyone who donates to certain organizations that in turn award scholarships to selected students.”141Id. In Espinoza, building on the newly invented anti-religious discrimination principle, the U.S. Supreme Court struck down this provision in the Montana Constitution.

Like Trinity Lutheran, it homed in on the status/use distinction to explain why the analogous Locke holding should not apply.142Id. at 2255–57. The Court emphasized that although both Espinoza and Locke addressed government scholarship funds used for religious education, the Montana Constitution prohibited all aid to sectarian schools simply by virtue of their being religious (that is, status) whereas the program in Locke excluded, specifically, just religious training (that is, use).143Id. at 2257. This case, the Court explained, “turns expressly on religious status and not religious use.”144Id. at 2256. It even took the time to refute claims that Montana’s Constitution was in fact about preventing “use” for religious education, responding that “[s]tatus-based discrimination remains status based even if one of its goals or effects is preventing religious organizations from putting aid to religious uses.”145Id. It asserted that “status-based discrimination is subject to ‘the strictest scrutiny.’ ”146Id. at 2257. Thus, a reasonable reading of the Court’s opinion would conclude that the status versus use distinction was central to this doctrine.

At the same time that it repeatedly emphasized its significance, however, the Court seemed to be readying itself to discard this distinction in the near future. It provided the caveat that “[n]one of this is meant to suggest that we agree . . . that some lesser degree of scrutiny applies to discrimination against religious uses of government aid.”147Id. Why then raise this distinction in the first place? As mentioned earlier, Trinity Lutheran was framed as a narrow decision addressing an even narrower, idiosyncratic, and relatively low-stakes set of facts. Allaying fears that it was anything broader than this, Trinity Lutheran’s footnote three had read: “This case involves express discrimination based on religious identity with respect to playground resurfacing. We do not address religious uses of funding or other forms of discrimination.”148Trinity Lutheran Church of Columbia, Inc. v. Comer, 582 U.S. 449, 465 n.3 (2017). Reliance on this status/use distinction, as well as the inclusion of this qualifying footnote, likely contributed to a majority that was able to bring along two justices (Breyer and Kagan) who shortly thereafter would pull away, dissenting in Espinoza and Carson.

Once the critical break with the religion clause precedent was achieved, like Lucy and Charlie Brown, the Chief Justice quickly pulled that football. It turns out Trinity Lutheran was no minor decision at all. In Carson, decided two years after Espinoza, the Court was clear that it was in fact Locke that was the minor decision. Leaving little ambiguity, Roberts asserted that “Locke cannot be read beyond its narrow focus on vocational religious degrees.”149Carson v. Makin, 142 S. Ct. 1987, 2002 (2022). Thus, just a short five-year time span had passed between Trinity Lutheran—adopting the status/use device as a central means of justifying its jarring divergence from Locke—and Carson—effectively retracting it. The unfortunate implication is that the status/use distinction served merely as a short-term results-oriented expedient—the proverbial camel’s nose that could push its way, ever so slightly, under the tent—facilitating the Court’s radical transformation of the Religion Clauses. 

In Espinoza, the Court repeatedly stressed the completely inapposite, but rhetorically powerful pejorative conception of “discrimination” to justify its holding, explaining that the Constitution “condemns discrimination against religious schools and the families whose children attend them.”150Espinoza, 140 S. Ct. at 2262. But even more than Trinity Lutheran, both Espinoza and Carson address a species of governmental action that is inevitably, and necessarily, grounded in discrimination—the state’s choices about education. It is indeed difficult to imagine a more consequential sphere of government speech than the fine-grained discretion involved when a democratically elected government chooses the ideas, ideals, knowledge, and values to impart to future generations. No question, this is most apparent in the field of public education, where states and localities are in the position of determining every last detail of a curriculum. But, unless it is establishing an open public forum, there is no reason to believe that it is less relevant when a state decides which educational alternatives it will choose to subsidize, and which it will not. Such choices are a direct manifestation of the will of the people as exercised by their elected representatives.

Indeed, the only limitation on this foundational majoritarian precept that it is “the people” who decide (indirectly, through elections) on the substance of public education and private educational subsidies, is when it is overridden by the Constitution itself, which, of course, requires a supermajority to overrule.151See Epperson v. Arkansas, 393 U.S. 97, 107 (1968). And one of the most notable examples of this can be found in the requirement of religious discrimination—that religion is subject to differential treatment—in the First Amendment. This requirement of religious discrimination in public education is well established in Court precedent.

In Epperson v. Arkansas, the Court confirmed that the religion clauses carve out an exception to the general and broad discretion a state has over its schools’ curricula.152Id. at 104–05. Under Arkansas law, public schools were prohibited from “teach[ing] the theory or doctrine that mankind ascended or descended from a lower order of animals.”153Id. at 98–99. The clear motivation behind the law was to thwart teaching that conflicted with the biblical account of the origin of life.154Id. at 109. Although the Court expressed a general reluctance to involve the judiciary in questions of educational policy, it was unequivocal that “the First Amendment does not permit the State to require that teaching and learning must be tailored to the principles or prohibitions of any religious sect or dogma.”155Id. at 106. The Court reaffirmed this reading in the 1987 decision Edwards v. Aguillard.156Edwards v. Aguillard, 482 U.S. 578, 594 (1987). This well-established understanding of the religion clauses, that educational choices which are otherwise within the discretion of state and local government must be judicially curtailed due to their religious nature, was not just contradicted, but inverted by Espinoza and Carson. The problem with “Montana’s no-aid provision” explains the Espinoza majority, is that it “bars religious schools from public benefits solely because of the religious character of the schools.”157Espinoza v. Mont. Dep’t of Revenue, 140 S. Ct. 2246, 2255 (2020).

Indeed, not only is the Court converting a constitutional principle that has always required differential treatment of religion into an anti-religious-discrimination rule, but government inaction—not doing what it was formerly required not to do by a conventional reading of the First Amendment—is understood as potentially coercive. As the Court explains, “[t]he Free Exercise Clause protects against even ‘indirect coercion,’ and a State ‘punishe[s] the free exercise of religion’ by disqualifying the religious from government aid . . . .”158Id. at 2256. Roberts, in other words, is taking Scalia’s Locke dissent logic one step further: not providing a government benefit is not just a relative “burden” on religion, it is a coercive punishment. Government benefits are so alluring that Jefferson’s separation of church and state is itself unconstitutional. The wall of separation is coercive because the church on one side will see the bag of goodies on the other side and feel compelled to un-church itself––to shed its religious identity so it too can get a hold of those benefits.

VIII.  THE ASYMMETRIC AND INTERDEPENDENT RELIGION CLAUSES

The Alice in Wonderland feel of the Court’s logic may be dizzying. But it is the built-in tension between the two religion clauses that makes the Court’s startling logical backflips possible. The Court is effectively borrowing concepts culled from one side of its religion clause decisions and lending them to the other. Since the two clauses were designed to pull in two different directions and operate in fundamentally different ways, predictably, the results are perverse.

While “neutrality” is drawn from Everson, “coercion” can be found in decisions such as 1992’s Lee v. Weisman. In that case, a student made an Establishment Clause challenge to a public school practice of inviting clergy members to give nondenominational prayers at graduation ceremonies. Although a passionate concurrence by Justices Blackmun, Stevens, and O’Connor argued for a more robust separationist rationale, the majority nonetheless struck down the policy, asserting that “at a minimum, the Constitution guarantees that government may not coerce anyone to support or participate in religion or its exercise.”159Lee v. Weisman, 505 U.S. 577, 587 (1992). Students, in other words, would feel peer pressure to conform to, and perhaps participate in, the religious exercise. This anti-coercion principle was firmly rooted in the Court’s Establishment Clause jurisprudence; the Court’s sights were set on identifying those types of government actions that cross the unconstitutional line of “respecting an establishment of religion.”

The Lee Court acknowledged that attendance at the ceremony was technically voluntary, but in the eyes of most students, it was a crucial rite of passage.160Id. at 594–95. The Court explained that “[i]t is a tenet of the First Amendment that the State cannot require one of its citizens to forfeit his or her rights and benefits as the price of resisting conformance to state-sponsored religious practice.”161Id. at 596. The Court in Espinoza and Carson takes this Establishment Clause principle, and applies it as if it were about free exercise. This is a mistake. These two clauses may work in tandem, but they function differently, as their disparate textual construction clearly suggests. The latter simply prevents the government from actively interfering with or “prohibiting” religious practice, whereas the former involves the thornier question of what it may mean for a law to “respect” an establishment of religion. As constitutional historian Leonard Levy explains, “Congress can pass laws regulating and even abridging the free exercise of religion without prohibiting it altogether.”162Levy, supra note 7. And not only does the Court, with little theoretical justification, blithely transfer an Establishment Clause test to a free exercise issue, it quietly alters its relative rigor.

As this concept of “coercion” is understood to be ever more capacious on the Free Exercise side of the ledger, including the “indirect” coercion of merely not having one’s religiously informed policy preferences fulfilled, the meaning of Establishment Clause coercion gets appreciably narrower. In Kennedy v. Bremerton School District, decided just one week after Carson, the Court appeared untroubled by establishment concerns because there was “no evidence” that, during a public prayer by an influential school employee at a public school event, “students [were] directly coerced to pray with [the coach].”163Kennedy v. Bremerton Sch. Dist. 142 S. Ct. 2407, 2419 (2022) (emphasis added). Thus, in the free exercise context, it would appear that a highly tenuous, and certainly debatable “indirect” form of coercion is sufficient to impose a constitutional demand that taxpayer money be used to fund private religion. At the same time, a popular football coach publicly praying “under the bright lights” of a stadium full of spectators,164Id. at 2439 (Sotomayor, J., dissenting). while “on duty,”165Id. at 2437 (Sotomayor, J., dissenting). and implicitly inviting student participation, was not a “direct” enough form of coercion to constitute an Establishment Clause violation. This coach had “made multiple media appearances to publicize his plans to pray at the 50-yard line,”166Id. at 2437 (Sotomayor, J., dissenting). and was someone from whom students might naturally seek favorable treatment such as extra playing time and recommendation letters167Id. at 2443 (Sotomayor, J., dissenting).. . Justice Sotomayor, in dissent, characterizes this newly watered down establishment test as “a nearly toothless version of the coercion analysis.”168Id. at 2434 (Sotomayor, J., dissenting). The effect is to invert the very meaning of the religion clauses, taking what would have been an unconstitutional violation of the Establishment Clause under the Court’s precedents—the injection of religion into the public schools—and transforming it into a constitutional requirement under the Free Exercise Clause.169Id. at 2441 (Sotomayor, J., dissenting).

Considering the ubiquity of both law and religion, and the fact that most policy will interact with religion in a multitude of ways, the task of drawing the establishment line is arguably much more difficult and subtle than drawing the free exercise line. On its face, the text of the Free Exercise Clause—a simple ban on governments prohibiting the free exercise of religion—would not seem to support a reading that demands active promotion by government of religion to preempt indirect coercion of religious believers who might feel left out. On its face, the Free Exercise Clause requires answering just two questions: First, how is a particular religion practiced, or exercised? Second, does the law at issue in fact prohibit that religion or its individual practitioners from practicing in such manner? The text of the Establishment Clause, in contrast, suggests that any state activity associated with, or part of a regime of, government establishment, should be subject to judicial scrutiny. The word “respecting” gives the Establishment Clause a degree of play that that the word “prohibiting” in the Free Exercise Clause does not.

As with all constitutional language, textual analysis allows for a range of plausible interpretations; the meaning given to both the word “prohibiting” and “respecting” is not fixed and will naturally be context dependent. As Randy Barnett explains, “[a]lthough most words are potentially vague, we do not face a problem of vagueness until a word needs to be applied to an object that may or may not fall within its penumbra.”170Randy E. Barnett, Interpretation and Construction, 34 Harv. J.L. & Pub. Pol’y 65, 68–69 (2011). The Janus-faced nature of the religion clauses—pushing in two different directions at the same time—heightens the interpretive challenge. Any doctrinal test by the Court that attempts to put flesh on the bones of the purportedly vague language in one religion clause, what Barnett refers to as a process of constitutional “construction,”171Id. at 69. must remain cognizant of its potential interaction with, impact on, or inconsistency with, the other clause. The Court’s insight of a “play in the joints”—a necessary degree of governmental discretion in enacting policies that promote the principles of one clause without violating the other—is consistent with this penumbral overlap.

Nonetheless, under many factual circumstances the same test simply cannot apply simultaneously under both the Establishment Clause and Free Exercise Clause without producing irreconcilable outcomes. The coercion test, so casually transferred from establishment to free exercise in Espinoza provides an example. The prayer in Lee is a violation of the Establishment Clause’s anti-coercion principle, but under the logic of Espinoza a constitutionally repaired, prayer-free graduation ceremony would be unconstitutionally coercive to religious students under the Free Exercise Clause by depriving them of a government benefit available to secular students. A free exercise anti-coercion rule would suggest that due to this deprivation, religious students would be indirectly coerced to either give up the benefit of publicly funded education and pay to attend a private religious school or relinquish their ability to partake in a religious graduation ceremony.

James Madison emphasized the importance of separation for the good of both government and religion, seeing it as a way of “[guarding against a] tendency to a usurpation on one side or the other, or to a corrupting coalition or alliance between them . . . .”172Geoffrey R. Stone, Louis Michael Seidman, Cass R. Sunstein, Mark V. Tushnet & Pamela S. Karlan, Constitutional Law 1438 (8th ed. 2018) (quoting James Madison). Roger Williams focused primarily on the way separation protects the church from control by the state.173Id. Yet, the majority in Espinoza dismisses this concern in just a few short paragraphs. Inverting historical reality, it treats Montana’s claim that “the no-aid provision promotes religious freedom”174Espinoza v. Mont. Dep’t of Revenue, 140 S. Ct. 2246, 2261 (2020). as the novel view, and its own recent invention of the anti-discrimination religion clauses as the constitutional baseline.

Consistent with an understanding that extends back hundreds of years, the state argued that “the no-aid provision protects the religious liberty of taxpayers by ensuring that their taxes are not directed to religious organizations, and it safeguards the freedom of religious organizations by keeping the government out of their operations.”175Id. at 2260. As if this deeply-rooted Madisonian understanding were a fringe perspective, the Court dismissed allowing an “infringement of First Amendment rights” on the basis of what it characterized as “a State’s alternative view.”176Id. But this is no “alternative view.” The dangers of the politicization of religion, the resentments taxpayer funding of religious institutions may engender, and the pressure governmental oversight and regulation would naturally place on the church, were not lost on the founders.

Effectively dismissing this wisdom in a single paragraph, the Court justifies its decision by emphasizing how its prior cases have allowed programs that provide aid to religious organizations where “attenuated by private choices.”177Id. at 2261. It then goes on to conflate freedom from government interference—these “private choices” that are rightfully protected under the Free Exercise Clause—and a right to non-discriminatory government benefits—which is, to the contrary, in direct tension with a traditional understanding and reading of the religion clauses. It achieves this slight-of-hand by citing for support its precedents that have “long recognized the rights of parents to direct ‘the religious upbringing’ of their children.”178Id. Of course, the freedom to opt-out of a majoritarian government program never implied a right to demand that the government offer an alternative version of that program that is tailored to one’s particular tastes.

Yet, in Carson v. Makin this is precisely what the Court requires of the state of Maine. The program at issue there, as discussed previously, differed from Espinoza in that it had limited its applicability based on the substance of the educational content of a school rather than its religious status. The Maine tuition assistance program was available to parents wishing to send their children to private schools in sparsely populated areas of the state where local government does not operate its own secondary school. Funds were ineligible however, if the desired school “promotes a particular faith and presents academic material through . . . that faith.”179Carson v. Makin, 142 S. Ct. 1987, 2001 (2022). The state explained that the private school option was designed to offer a “rough equivalent” of the secular public schools available in more populous parts of the state. The Carson family, however, wanted to send their daughter to a private school with a “Christian worldview [that] aligns with their sincerely held religious beliefs.”180Id. at 1994. Under the Court’s new anti-religious discrimination reading, the state was now required to use taxpayer funds to accommodate the family’s religious tastes.

IX.  THE DEMISE OF THE STATUS/USE DISTINCTION

Unless a majoritarian democracy is structured to require unanimity, it is inescapable that some minority of the population will be unhappy with the substantive policy choices the government makes. As Alexander Tsesis has pointed out, “[there are] disagreements about the wisdom of myriad government programs, policies, statutes, and priorities.”181Alexander Tsesis, Government Speech and the Establishment Clause, 2022 U. Ill. L. Rev. 1761, 1771 (2022). A distinct policy choice to fund only private schools with an evidence-based curriculum, is, of course, bound to displease those who prefer a faith-based approach to education. However, a state may have many legitimate policy reasons for declining to fund religious education, and these reasons may be independent of a desire to adhere to a “stricter separation of church and state than the Federal Constitution requires.”182Carson, 142 S. Ct. at 1997. Most obviously, a government may conclude that an epistemological approach grounded in faith is in tension with a commitment to the scientific method. Its reasoning, in other words, may relate directly to its judgment as to how it will best fulfill its educational mission. The Court acknowledges that only private schools that “meet certain basic requirements” were eligible to receive the funds under Maine’s program.183Id. at 1993. Yet, somehow, four pages later, the Court characterizes it as “a neutral benefit program,” seemingly forgetting that the state established detailed criteria laying out just what attributes schools must have if it is to fund them.184Id. at 1997.

With Carson, the Court thus ratchets up its novel anti-religious-discrimination interpretation of the religion clauses to include substantive as well as status-based distinctions. As the Court explains, “the prohibition on status-based discrimination under the Free Exercise Clause is not a permission to engage in use-based discrimination.”185Id. at 2001. The former was at least arguably one step further removed from the kind of policy discretion essential for responsive democratic judgment––a discretion that informs the Court’s own government speech doctrine. In theory, status-based distinctions are also potentially indicative of a substance-free animus or discriminatory impulse against religion. But “use-based discrimination,” as the Court puts it, is just ordinary lawmaking. As preeminent constitutional historian Leonard Levy unequivocally concluded, “the fact is that no framer believed that the United States had or should have power to legislate on the subject of religion.”186Levy, supra note 7, at 121–22. Yet, perversely, under the Court’s new anti-religious-discrimination doctrine, states now must do so. As of 2022, the substantive educational content a state chooses not to expend its resources on is subject to the Court’s intrusive new religion clause rule.

Although those who want their children to receive a faith-based education are by no means precluded from making this choice, according to the Court the mere fact that they must pay for such education themselves (while the choice to utilize a secular private school would be supported by the state) exerts coercive pressure on their choice.187Carson, 142 S. Ct. at 1996. A failure to fund faith-based approaches to education does not just result in the natural disappointment felt by those in a democracy whose policy preferences do not go completely fulfilled, such failure to spend “ ‘penalizes the free exercise’ of religion.”188Id. at 1997. The implications of this conceptualization are quite stunning. The Supreme Court is effectively depriving democratic governments of their discretion to determine their spending priorities in one of the most consequential and democratically hard-fought domains: public education.

X.  THE LIMITING PRINCIPLE PROBLEM

Now, some might be inclined to see the concerns above as alarmist. Carson, after all, addresses just one case-specific state program. However, it is difficult to see the stopping point of the Court’s logic. The Court’s novel anti-religious discrimination rule lacks a limiting principle. In Trinity Lutheran, Roberts seemed at least mildly attuned to this potential concern by emphasizing the purportedly status-based nature of the discrimination. But, consistent with the Chief’s camel’s-nose-under-the-tent approach to doctrinal change, after Carson, any government program might become the next target of an allegation that it is discriminating against religion, and therefore violating the Free Exercise Clause. Under the Court’s newly expansive anti-religious-discrimination rule in Carson, simply not offering a comparable religion-based alternative to any secular state benefit presents a potential constitutional infraction. What might the future portend under such a regime? We might anticipate a kind of constitutionally mandated menu-based governance in which state resources must be shared equally among religious and non-religious options.

For religion, this vision may ultimately prove to be self-defeating. Government resources are limited, as is the tolerance of the populace for ever higher taxes. Constitutionally mandated religious alternatives will become costly and will ultimately be subjected to the same type of politicized, compromise-laden, and messy process that is at the heart of all spending decisions in a democratic polity. As a constitutionally imposed unfunded mandate, religion would lose its prized independence.

Granted, the anti-religious-discrimination impulse is understandable. As mentioned, the new anti-religious-discrimination Free Exercise principle is no doubt rooted to some extent in a broader concern that religion, religious belief, and religious practitioners have been unfairly mistreated and disparaged by a secular society. However, those who would like to see an expanded role for religion in the public sphere, even those who support taxpayer subsidies of religion in certain areas, may ultimately find themselves deeply troubled by the ultimate consequences of the slippery-slope the Court has erected. The Court’s radical re-interpretation of the religion clauses may prove self-defeating, for government and religion. Without a limiting principle, it cannot be contained.

XI.  THE PUBLIC FORUM DOCTRINE TO THE RESCUE

All of this is not to say that there is no place for a constitutional principle prohibiting, in some contexts, discrimination against religion. Just because the religion clauses demand the opposite, does not mean there are not other settings in which religion may be protected from government. The Free Exercise Clause, most obviously, protects religion by forbidding targeted prohibitions on free exercise. But those who would like to see a greater presence of religion in the public sphere have an alternative constitutional hook to grasp. Another First Amendment doctrine, derived from the Free Speech Clause, does include an anti-discrimination rule that serves to protect against forms of religious discrimination.

The public forum doctrine prohibits the government from imposing viewpoints, and sometimes content-based, discrimination on private speech; and the Court has concluded that this restriction extends to religious expression.189See Lamb’s Chapel v. Ctr. Moriches Union Free Sch. Dist., 508 U.S. 384 (1993); Rosenberger v. Rector & Visitors of Univ. of Va., 515 U.S. 819 (1995). The Court reminded us most recently of this principle in Shurtleff v. City of Boston. The case involved a government program that over time allowed hundreds of private groups to fly their flags outside of Boston’s city hall. The city, however, denied such opportunity to a Christian group. In ruling against Boston on free speech grounds, the Court explained that “[w]hen a government does not speak for itself, it may not exclude speech based on ‘religious viewpoint’; doing so ‘constitutes impermissible viewpoint discrimination.’ ”190Shurtleff v. City of Boston, 142 S. Ct. 1583, 1593 (2022).

If the government speech doctrine can be said to be pro-discrimination—rooted in the understanding that a democratically accountable government must have the ability to be selective as to what policy messages it will, or will not, send—its cousin, the public forum doctrine, forbids discrimination in government-owned, funded, or controlled forums. Once a government opens property or a program up to the broader public, establishing a public forum—or to a select portion of the public for more circumscribed purposes, establishing what the Court has called a “limited” public forum—it may not discriminate on the basis of “content” (or merely “viewpoint” where the public forum is limited).191See, e.g., McCullen v. Coakley, 573 U.S. 464 (2014); Christian Legal Soc’y v. Martinez, 561 U.S. 661 (2010). Government speech and public fora may be conceived as two poles on opposite ends of a single continuum, with government having almost complete control over what is or is not expressed on the government speech end and minimal power to restrict or dictate expression on the other.192See Wayne Batchis, The Government Speech-Forum Continuum: A New First Amendment Paradigm and Its Application to Academic Freedom, 75 N.Y.U. Ann. Surv. Am. L. 33 (2019). The critical point is that the public forum doctrine, unlike the Court’s new anti-religious-discrimination rule, has a clear limiting principle: a government program must fall within the definition of a public forum (or limited public forum) for religion to receive protection from discrimination. Otherwise, a policy choice, and any messages associated with it—unless, of course, it “prohibits” or “respects an establishment of religion”—would be treated as government speech.

Thus Justice Thomas, in his Espinoza concurrence, begs the question when he criticizes the “strict separation” approach to the religion clauses for the way it would ostensibly remove “the entire subject of religion from the realm of permissible governmental activity . . . operat[ing] as a type of content-based restriction . . . .”193Espinoza v. Mont. Dep’t of Revenue, 140 S. Ct. 2246, 2266 (2020) (Thomas, J., concurring). Religion would not be banished from the public sphere under the traditional, straight-forward reading of the First Amendment advocated in this article; the impact of the Religion Clauses would simply turn on whether or not the government itself is speaking or whether its “activity” was creating a public forum. As the Court has repeatedly reaffirmed, government speech is all about content-based restrictions on speech—both the discriminating choices government makes as to what messages it will or will not devote its resources to, and the structural boundaries enshrined in the Constitution that may similarly shape, limit, or direct its expressive choices. A public forum, on the other hand, does demand that the government avoid content or viewpoint-based discrimination.

Indeed, this is where a misguided concurrence by Justice Kavanaugh in Shurtleff gets it so wrong. He seeks to supplement the majority’s opinion by emphasizing that a government does not merely violate the Establishment Clause by treating religion equally to other government beneficiaries. Equal (favorable) treatment of religion by government is permissible under certain circumstances in accordance with both the public forum doctrine and the “play in the joints” religion clause principle long accepted by the Court.194See supra note 61and accompanying text. But in startlingly broad terms, Kavanaugh goes on to assert that “a government violates the Constitution when (as here) it excludes religious persons, organizations, or speech because of religion from public programs, benefits, facilities, and the like.”195Shurtleff, 142 S. Ct. at 1594 (Kavanaugh, J., concurring). Confined to public fora, such a statement of the rule may be true; but outside of these confines, such a rule would impose constitutionally illimitable unfunded expressive mandates on governments, potentially violating anti-establishment principles and the core premise of the government speech doctrine along the way.  

In contrast, drawing a boundary between a public forum—where religious expression would be protected from discrimination—and government speech—where government would have the option and sometimes obligation to discriminate against religious messages––is remarkably consistent with the inherent tension built into the Religion Clauses. It brings the First Amendment full circle, connecting the Speech and Religion Clauses in a logically coherent way. Governments may establish public fora to facilitate private speech, as governments have a rich and important history of doing, whether it is a public park, the after-hours use of public facilities for associational meetings, or a public university’s student organization program. These are venues that may be owned and maintained by the state, but as public forums, the speech that occurs there is protected and not presumed to represent the government’s voice. As a result, such expression, even if overtly religious, is unlikely to raise traditional Establishment Clause concerns; it is unlikely to generate the impression of government endorsement or to have a coercive effect. Government, in other words, would be free to facilitate free exercise values in a way that is cognizant of Establishment Clause values, while at the same time acting consistently with free speech doctrine.

CONCLUSION

The Supreme Court’s decision in Carson v. Makin is the third in a trilogy of cases dramatically upending the meaning of the First Amendment’s Religion Clauses. Beginning with Trinity Lutheran in 2017, and followed by Espinoza in 2020, the Court has moved forward with an aggressive project of transforming the Religion Clauses into a broad anti-religious-discrimination clause. In this paper, I traced this doctrinal devolution and argued that the Court’s novel reinterpretation is deeply misguided.

By design, the Religion Clauses require discrimination—religion is to be treated differently from non-religion in a broad range of state action. The Establishment Clause targets religion specifically by prohibiting laws that intermingle government with religion in an impermissible manner—whereas intermingling government with other philosophies, worldviews, institutions, or sets of values is a perfectly ordinary and generally acceptable aspect of policymaking. The Free Exercise Clause likewise forbids government interference with religious practice—whereas government is certainly free to, and is indeed expected to, interfere with a vast range of non-religion related conduct deemed to violate criminal and civil law. Religion, in short, is different. The contemporary Supreme Court, however, has inverted this most basic insight.

The Court’s new Religion Clause jurisprudence is also on a collision course with its burgeoning government speech doctrine. That doctrine recognizes that in a democratic polity, every policy choice entails paths not chosen. Government must be able to select its own message, and in turn, discriminate against those messages it wishes not to communicate. While there are some exceptions to the rule—specifically, the boundaries set by the Constitution itself—the default is governmental discretion, tempered only by accountability at the ballot box. Thus, the Religion Clauses, in conjunction with the government speech doctrine, mandate that government either be free to speak with its own voice when it is acting within the “play in the joints” in between the two clauses, or treat religion distinctly—to discriminate—when required to do so under the Constitutional mandate of establishment or free exercise.

To say that discrimination is required under the Free Exercise or Establishment Clause is not to say discrimination against religion is always constitutional. Outside of the Religion Clauses, other protections against objectionable discrimination remain. The Court’s public forum doctrine, for example, protects free expression of religion from content-based discrimination when the government itself is not speaking. Adverse or favored treatment by government targeting religion generally, particular religious sects, or particular religious practices, may be impermissible. But when it comes to the Religion Clauses, these are circumstances in which the discrimination provides evidence that the government is either prohibiting free exercise or making a law respecting an establishment of religion. The discrimination itself is not the Constitutional offense. Acting as if it is, is highly misleading. The Religion Clauses provide nothing like the broad anti-discrimination mandate today’s Court imputes to them. They demand the opposite.

The heart of the Court’s recent trilogy of cases—from Trinity Lutheran v. Comer to Carson v. Makin—is a constitutional mandate that government subsidize religious speech to avoid a Religion Clause “discrimination” claim. It is a command that government express ideas it may not wish to express. The Court’s reimagining of its Religion Clauses jurisprudence is inconsistent with the First Amendment’s original meaning, anti-democratic, and in direct tension with the government speech doctrine.

97 S. Cal. L. Rev. 367

Download

* J.D., PhD.; Professor and Director of Legal Studies, University of Delaware, Department of Political Science and International Relations.

Regressive White-Collar Crime

Fraud is one of the most prosecuted crimes in the United States, yet scholarly and journalistic discourse about fraud and other financial crimes tends to focus on the absence of so-called “white-collar” prosecutions against wealthy executives. This Article complicates that familiar narrative. It contains the first nationwide account of how the United States actually prosecutes financial crime. It shows—contrary to dominant academic and public discourse—that the government prosecutes an enormous number of people for financial crimes and that these prosecutions disproportionately involve the least advantaged U.S. residents accused of low-level offenses. This empirical account directly contradicts the aspiration advanced by the FBI and Department of Justice that federal prosecution ought to be reserved for only the most egregious and sophisticated financial crimes. This Articles argues, in other words, that the term “white-collar crime” is a misnomer.

To build this empirical foundation, the Article uses comprehensive data of the roughly two million federal criminal cases prosecuted over the last three decades matched to county-level population data from the U.S. Census. It demonstrates the history, geography, and inequality that characterize federal financial crime cases, which include myriad crimes such as identity theft, mail and wire fraud, public benefits fraud, and tax fraud, to name just a few. It shows that financial crime defendants are disproportionately low-income and Black, and that this overrepresentation is not only a nationwide pattern, but also a pattern in nearly every federal district in the United States. What’s more, the financial crimes prosecuted against these overrepresented defendants are on average the least serious. This Article ends by exploring how formal law and policy, structural incentives, and individual biases could easily create a prosecutorial regime for financial crime that reinforces inequality based on race, gender, and wealth.

INTRODUCTION

Fraud is an old crime. It can be found in criminal codes around the world for as long as the historical record exists. The Code of Hammurabi, composed around 1750 B.C.E. in Ancient Babylon, included several provisions outlawing various forms of fraud with punishments including death.1Martha T. Roth, Laws of Hammurabi, in Law Collections from Mesopotamia and Asia Minor 82–84, 105, 130 ¶¶ 9, 11, 126, 265 (Piotr Michalowski ed., 1997). As Alice Ristroph has noted, the second-lowest level of Hell in Dante’s fourteenth-century Inferno is reserved for people who perpetrate fraud, treating them more harshly than those who engage in physical violence.2Alice Ristroph, Criminal Law in the Shadow of Violence, 62 Ala. L. Rev. 571, 620–21 (2011) (quoting Dante, The Inferno, Canto XI, 23–29). On the other hand, fraud and financial crimes are capable of causing significant physical harm and, for that reason, some resist labeling white-collar crime as “non-violent.” See, e.g., Miriam Saxon, Subcomm. on Crime of the Comm. on the Judiciary, 95th Cong., 2d. Sess., White Collar Crime: The Problem and the Federal Response 4 (Comm. Print 1978) (“[P]articularly in those many instances of economic crime in which hundreds or thousands of people are affected, the harm to society can frequently be described as violent.”). According to the United States Supreme Court, “fraud has consistently been regarded as such a contaminating component in any crime that American courts have, without exception, included such crimes within the scope of moral turpitude.”3Jordan v. De George, 341 U.S. 223, 229 (1951).

Our state­ and federal criminal codes define myriad kinds of frauds, which comprise the majority of what we call “financial crimes” or “white-collar crimes.” Every year, tens of thousands of U.S. residents are convicted of financial crimes, most of them frauds.4See infra Section I.B.1. Yet, financial crime rarely surfaces in public discussion about how substantive criminal law fuels mass and unequal incarceration in the United States.

Instead, the terms “financial crime” and “white-collar crime” usually conjure up images of a rich banker on Wall Street or an elite executive in a powerful multinational corporation who is able to escape prosecution.5See Samuel W. Buell, “White Collar” Crimes, in The Oxford Handbook of Criminal Law 839 (Markus D. Dubber & Tatjana Hörnle eds., 2014) (noting public discussion of white-collar crime tracks a definition that includes bankers, “Wall Street,” or “corporate America,” as well as “professionals and other service providers and gatekeepers, such as lawyer and accountants, who are integral to the corporate world”). This imagery is fueled by an academic and popular discourse that tends to equate financial crime with the executive class and emphasizes the absence of prosecutions against wealthy people. For example, in recent years, much journalistic coverage of financial crime has focused on explaining why so few people and no companies were convicted of a crime connected to the financial crisis of 2008.6See, e.g., Miriam Baer, Myths and Misunderstandings in White-Collar Crime 108 (2023) (“Commentators simply cannot fathom why federal prosecutors were unable to mount cases against the architects of the subprime crisis, a crisis that is commonly described as one big scam.”); Jesse Eisinger, The Chickenshit Club: Why the Justice Department Fails to Prosecute Executives (2017); Patrick Radden Keefe, Why Corrupt Bankers Avoid Jail, New Yorker (July 31, 2017), https://www.newyorker.com/magazine/2017/07/31/why-corrupt-bankers-avoid-jail [https://perma.cc/QU8K-2UUX]; Michael Winston, Why Have No CEOs Been Punished for the Financial Crisis?, The Hill (Dec. 8, 2016, 6:10 PM), https://thehill.com/blogs/pundits-blog/finance/309544-why-have-no-ceos-been-punished-for-the-financial-crisis [https://perma.cc/4SWK-HFGA]; William D. Cohan, A Clue to the Scarcity of Financial Crisis Prosecutions, N.Y. Times (July 21, 2016), https://www.nytimes.com/2016/07/22/business/dealbook/a-clue-to-the-scarcity-of-financial-crisis-prosecutions.html? [https://perma.cc/8DTA-JKK3]; William D. Cohan, How Wall Street’s Bankers Stayed Out of Jail, Atlantic (Sept. 2015), https://www.theatlantic.com/magazine/archive/2015/09/how-wall-streets-bankers-stayed-out-of-jail/399368 [https://perma.cc/P3BD-7BYU]. Similarly, much academic scholarship about financial crime attempts to document and explain the causes and consequences of the U.S. Department of Justice’s (“DOJ”) routine policy of declining or deferring prosecution of financial crimes committed by or within large companies.7See, e.g., W. Robert Thomas, Incapacitating Criminal Corporations, 72 Vand. L. Rev. 905 (2019); Nick Werle, Note, Prosecuting Corporate Crime When Firms Are Too Big to Jail: Investigation, Deterrence, and Judicial Review, 128 Yale L. J. 1366 (2019); Mihailis E. Diamantis, Clockwork Corporations: A Character Theory of Corporate Punishment, 103 Iowa L. Rev. 507 (2018); Brandon L. Garrett, Too Big to Jail: How Prosecutors Compromise with Corporations (2014); Jennifer Arlen, Prosecuting Beyond the Rule of Law: Corporate Mandates Imposed Through Deferred Prosecution Agreements, 8 J. Legal Analysis 191 (2016); Jennifer Arlen & Marcel Kahan, Corporate Governance Regulation Through Nonprosecution, 84 U. Chi. L. Rev. 323 (2017). But see Samuel W. Buell, Is the White Collar Offender Privileged?, 63 Duke L. J. 823, 824–25 (2014) (questioning the validity of the popular belief that the American criminal system favors corporate offenders). This Article argues that this popular conception of financial crime is inaccurate.

The popular imagery surrounding white-collar crime is also kindled by prosecutors themselves. For decades, the DOJ has repeatedly and publicly touted its focus on fraud prosecutions that hold corporate executives and corporations accountable as opposed to poor and middle-class people. Prosecuting business executives, according to Attorney General Merrick Garland, is “essential to Americans’ trust in the rule of law.”8Attorney General Merrick B. Garland, Remarks to the ABA Institute on White Collar Crime (Mar. 3, 2022) (transcript available at https://www.justice.gov/opa/speech/attorney-general-merrick-b-garland-delivers-remarks-aba-institute-white-collar-crime [https://perma.cc/V9MT-8FU3]). That is because “the rule of law requires that there not be one rule for the powerful and another for the powerless; one rule for the rich and another for the poor.”9Id. Numerous attorneys general have made similar statements.10For example, in 2002, Attorney General John Ashcroft compared corporate fraud with the September 11th attacks. While those attacks were an assault on freedom “from abroad,” corporate fraud, according to Ashcroft, was an assault on freedom “from within.” Attorney General John Ashcroft, Enforcing the Law, Restoring Trust, Defending Freedom, Remarks to the Corporate Fraud/Responsibility Conference (Sept. 27, 2002) (transcript of remarks as prepared at https://www.justice.gov/archive/ag/speeches/2002/092702agremarkscorporatefraudconference.htm [https://perma.cc/P5RZ-SBT2]). In a 2014 speech about corporate crime, Attorney General Eric Holder boasted that DOJ charged more white-collar defendants between 2009 and 2013 than during any previous five-year period going back to at least 1994. Attorney General Eric Holder, Remarks on Financial Fraud Prosecutions at NYU School of Law (Sept. 17, 2014) (transcript available at https://www.justice.gov/opa/speech/attorney-general-holder-remarks-financial-fraud-prosecutions-nyu-school-law [https://perma.cc/HA2A-QFBK]). This prosecutorial discourse risks creating the false impression that financial crime is primarily committed by the most wealthy and privileged Americans and, perhaps as a result, is leniently, if ever, punished. The reality, as this Article shows, is the opposite.

In short, this Article shows that our prevailing conception of financial crime is, at best, incomplete and, at worst, wrong. It argues that scholarly and public discourse around financial crime, which focuses on the absence of “white-collar” prosecutions (that is, prosecutions of members of the wealthy executive class), paints an inaccurate picture of how financial crime is prosecuted. The United States does, in fact, prosecute a huge number of people for financial crimes—thousands per year. But these defendants are for the most part not wealthy executives. Instead, financial crime prosecutions disproportionately involve people who are low-income and people who are Black. This Article suggests that financial crime is in this way unexceptional in an American criminal system that otherwise consistently reflects class- and race-based inequality.11This Article thus suggests that the notion “carceral exceptionalism” in the context of white-collar crime is misguided. See Benjamin Levin, Mens Rea Reform and Its Discontents, 109 J. Crim. L. & Criminology 491, 548-57 (2019) (identifying “carceral exceptionalism” as the phenomenon in which “scholars and advocates on the left” favor “the full force of the carceral state” for certain “exceptional” defendants).

With data on the roughly two million federal criminal cases prosecuted since the early 1990s matched with county-level Census data, this Article is the first comprehensive study of all federal financial crime prosecutions.12As described in Section I.C, others explored similar questions in a series of studies produced in the 1980s through early 2000s known as the “Yale Studies.” The Yale Studies focused on 210 white-collar defendants prosecuted in seven federal district courts. See infra notes 112–116 and accompanying text. This Article demonstrates that, like all federal criminal defendants, the people convicted of financial crimes have fewer resources than the average U.S. adult. Financial crime defendants have attained less formal education than average and frequently rely on appointed counsel. Federal judges waive the fines of roughly eighty-six percent of federal financial crime defendants due to the defendant’s inability to pay. In other words, the median fine in a federal white-collar prosecution is $0.

This Article also shows that financial crimes are not prosecuted at equal rates across the U.S. population. Women are prosecuted at higher rates for financial crimes than for other types of federal crimes.13See infra Appendix Table A.3 (noting that women make up roughly thirty percent of federal financial crime defendants and roughly fourteen percent of all federal criminal defendants). The same is true in state courts, as Kaaryn Gustafson and others have pointed out.14See Kaaryn S. Gustafson, Cheating Welfare: Public Assistance and the Criminalization of Poverty 7 (2011) (pointing out that prosecutions of fraud are “unusual” in that they are more frequently prosecuted against women than other types of crimes); see also Brian A. Reaves, U.S. Dep’t of Just., Bureau of Just. Stat., Felony Defendants in Large Urban Counties, 2009 – Statistical Tables 5 (2013), https://bjs.ojp.gov/content/pub/pdf/fdluc09.pdf [https://perma.cc/C74Y-LETR] (“In 2009, the most frequently charged offenses among female felony defendants were fraud (37%), forgery (34%), and larceny/theft (31%).”). Financial crime prosecutions are also unequal by race. Black women are especially likely to be prosecuted for financial crimes and are prosecuted at roughly three times the per capita rate as Hispanic and non-Hispanic White women.15See infra Appendix Table A.3. The same is true for Black men, who are prosecuted at roughly three times the rate as Hispanic and non-Hispanic White men.16Id.

This analysis is especially important for understanding racial inequalities among female defendants. Black women are more likely to be convicted of a financial crime than any other type of federal crime.17This observation is based on the author’s analysis of the data. The data used in this paper is available for download at Stephanie Holmes Didwania, Data for “Regressive White-Collar Crime, Nw. Univ. (2024), https://doi.org/10.21985/n2-gav7-wt94 [hereinafter Didwania, Data]. This has been true every year since 1994—as far back as reliable federal criminal case data goes.18Id. The same is not true of any other race-gender group of defendants.19Id. Like Black women, non-Hispanic White women and non-Hispanic women of another race are prosecuted for financial crimes more than any other type of crime. Unlike Black women, however, this has not been the case every year for women who are not Black.

This Article also shows it is not the case that the defendants most overrepresented in financial crime cases (that is, low-income defendants and Black defendants) commit the most severe or complex financial crimes. The opposite is true. I argue that these prosecutorial patterns could easily stem from a combination of formal law and policy, individual biases, and systemic incentives.

A muddled view of how financial crime is prosecuted has meaningful consequences. Maybe because financial crime (often stylized as “white-collar” crime) is viewed as a pursuit of the elite, there seems to be little appetite for leniency toward those convicted of financial crimes on either side of the political aisle. As Benjamin Levin and Kate Levine write, “prosecuting some imagined class of bankers or executives remains very popular with many liberal, left, and progressive commentators.”20Benjamin Levin & Kate Levine, Redistributing Justice, Colum. L. Rev. 26 (forthcoming 2024). See also Douglas Husak, The Price of Criminal Law Skepticism: Ten Functions of the Criminal Law, 23 New Crim. L. Rev. 27, 51-52 (2020) (“Even those members of the public who tend to agree that the criminal justice system punishes too many persons with too much severity can be heard to complain when leniency is afforded to . . . white collar criminals.”). Along these lines, President Biden’s clemency efforts have almost exclusively—and in some cases explicitly—focused on people serving sentences for drug trafficking or possession.21For example, in September 2021, the Biden administration invited federal prisoners to apply for clemency if they had been released home under the pandemic relief bill and had four years or less remaining on their sentences. The invitation was limited, however, to people who had been convicted of drug crimes. Sam Stein, Biden Starts Clemency Process for Inmates Released due to Covid Conditions, Politico (Sept. 13, 2021, 1:17 PM), https://www.politico.com/news/2021/09/13/biden-clemency-covid-inmates-511658 [https://perma.cc/N93A-GUEM]. In April 2022, Biden took his first formal clemency actions as President, granting three pardons and seventy-five commutations. Press Release, White House, Clemency Recipient List (Apr. 26, 2022), https://www.whitehouse.gov/briefing-room/statements-releases/2022/04/26/clemency-recipient-list [https://perma.cc/6RQG-NGMV]. Of the seventy-eight clemency recipients, all but one had been convicted of drug crimes. Id. In October 2022, Biden announced a pardon of all prior federal convictions of marijuana possession. Press Release, White House, Statement from President Biden on Marijuana Reform (Oct. 6, 2022), https://www.whitehouse.gov/briefing-room/statements-releases/2022/10/06/statement-from-president-biden-on-marijuana-reform [https://perma.cc/3X8W-U2UE].

Not only has the Biden administration essentially excluded white-collar prisoners from its clemency efforts, but Attorney General Merrick Garland has also emphasized that cracking down on white-collar crime is one of DOJ’s top priorities.22Garland, supra note 8. In a March 2022 speech describing this white-collar initiative, then-Assistant Attorney General for the Criminal Division Kenneth A. Polite, Jr. echoed the idea that white-collar crime is not punished harshly enough, telling the audience, “When we talk about drug dealing and violence, we all have no problem conjuring notions of accountability for the criminal actors. But the sheer mention of individual accountability in white-collar cases was, and is, received as a shockwave in our practice.”23Assistant Attorney General Kenneth A. Polite, Jr., Justice Department Keynote at the ABA Institute on White Collar Crime (Mar. 3, 2022) (transcript of remarks as prepared for delivery at https://www.justice.gov/opa/speech/assistant-attorney-general-kenneth-polite-jr-delivers-justice-department-keynote-aba [https://perma.cc/L8UU-FFBB]). This Article cautions that directing more resources toward prosecuting white-collar crime could perpetuate class- and race-based inequalities rather than mitigate them.24This Article thus supports the argument advanced by Benjamin Levin and Kate Levine that those on the progressive left who hope the criminal system will work as a tool of progressive redistribution is unlikely to succeed. Levin & Levine, supra note 20, at 37–38 (forthcoming 2024) (arguing that “institutions of the punitive state are inherently regressive and are antithetical to the egalitarian vision articulated by many of the commentators who have embraced redistributive carceral projects”)

The federal criminal system is a worthy site to study the regressive prosecution of white-collar crime even though most criminal defendants in the United States are prosecuted in state courts.25In 2020—the last year for which data was available—around 1.2 million people were under the legal jurisdiction of a state or federal correctional authority. Within this population, eighty-seven percent of the people were under state jurisdiction and thirteen percent were under federal jurisdiction. This calculation excludes people held in local jails. See E. Ann Carson, U.S. Dep’t of Just., Bureau of Just. Stat., Prisoners in 2020 – Statistical Tables 7 (2021). This Article focuses on the federal system for two reasons. First, the federal criminal system is important in its own right. The federal government incarcerates more people than any state and federal prisoners on average serve longer sentences than state prisoners.26Id. at 7–8 (showing that the federal prisoner population was 152,156 in 2020 and the jurisdiction with the second-largest prisoner population (Texas) imprisoned 135,906 people in 2020). The median time served in state prison for prisoners released in 2018 was 1.3 years. Danielle Kaeble, U.S. Dep’t of Just., Bureau of Just. Stat., Time Served in State Prison, 2018 1 (2021). By contrast, the median federal sentence in the 1994–2019 period was two years. Federal criminal defendants must serve at least eighty-five percent of their sentence, so even accounting for good time credit, the median time served for federal prisoners over this period was at least 1.7 years. 18 U.S.C. § 3624(b)(1) (providing that federal prisoners serving more than 1 year in prison can get credit towards their sentence of 54 days per year if they display “exemplary compliance with institutional disciplinary regulations”). Fraud—the most common financial crime—is itself the third-most prosecuted type of federal crime after drug trafficking and immigration offenses.27This observation is based on the author’s analysis of the data. See Didwania, Data, supra note 17. Indeed, even as federal prosecutions of other types of crimes have exploded, fraud alone has constituted around 10 percent or more of the federal felony docket since the early 1990s.28Id.

Second, as described in Section I.B, federal officials repeatedly emphasize that it is their goal to prosecute the most egregious and complex financial crimes. Because state courts have concurrent jurisdiction over many financial crimes, DOJ and FBI can in theory focus their efforts on complex investigations. DOJ and FBI routinely tout their partnerships with other federal agencies to detect and prosecute sophisticated financial crimes. It seems unlikely that state prosecutors are doing better than the federal government at prosecuting complex financial crimes with fewer investigative resources. For this reason, prosecuting serious financial crime is often viewed as a federal project.29See Daniel C. Richman & William J. Stuntz, Al Capone’s Revenge: An Essay on the Political Economy of Pretextual Prosecution, 105 Colum. L. Rev. 583, 601–02 (2005).

Indeed, many observers rightly view the complexity of serious financial crimes as an impediment to prosecution. Criminal investigations can take years; relevant documents can number in the millions; trials can take months.30See, e.g., Press Release, U.S. Dep’t. of Just., Federal Jury Convicts Former Enron Chief Executives Ken Lay, Jeff Skilling on Fraud, Conspiracy, and Related Charges (May 25, 2006), https://www.justice.gov/archive/opa/pr/2006/May/06_crm_328.html [https://perma.cc/9UYY-Y7LE] (noting that the trial of Enron executives Kenneth Lay and Jeffrey Skilling took fifty-six days). This Article’s primary goal is not to determine whether the federal government has chosen the best balance in prosecuting the cases that it does, but rather to bring to light the fact that most financial crime cases are modest ones that disproportionately impact people with the fewest advantages.

This Article’s analysis advances in four steps. Part I traces the history of financial crime and shows how, for centuries, rich and powerful people have escaped prosecution for financial crimes while people who are poor and middle-class have been prosecuted. Section I.B describes how federal financial crime cases are prosecuted today and provides examples of four such cases. Section I.C argues that most scholarly and public discourse around financial crime overlooks the types of financial criminal cases that are most routinely prosecuted in U.S. courts.

Part II presents the bulk of the empirical analysis. It shows persistent income, gender, and race gaps in financial crime prosecutions that disfavor defendants who are low-income, male, and Black. Part III offers many possible explanations for the results. It groups these explanations into four categories. First, Section III.A considers but rules out the possibility that people who are overrepresented commit the most serious financial crimes. Second, Section III.B describes how systemic and structural conditions create a system in which prosecutors are motivated to prosecute the cases they view as most winnable. Third, Section III.C describes ways that formal criminal law and policies could lead prosecutors to focus their efforts on simplistic, low-level financial crimes. As one example, it shows how federal laws governing restitution benefit defendants with more resources. Finally, Section III.D describes how biases on the part of actors in the criminal system could contribute to inequality.

Part IV concludes. It argues that the findings provide vital context for understanding how financial crime is prosecuted in the United States and challenges the popular notion that financial crime is under-prosecuted.

I.  PROSECUTING FINANCIAL CRIME

This Part broadly traces the history of financial crime prosecution. As described in Section I.A, the United States has a long history of prosecuting poor and middle-class people for financial crimes. (Part II shows that this pattern continues through today, despite repeated statements to the contrary by modern prosecutors). Section I.B describes how the federal government has prosecuted fraud since the 1990s and presents four archetypical examples of federal financial crime cases, to which I return throughout the Article. Section I.C explains how this Article contributes to the existing literature on federal financial crime, which largely avoids discussing the relatively low-level cases that pervade the federal criminal system.

A.  Early Prosecutions and the Concept of “White-Collar” Crime

Most financial crimes are frauds.31Other financial crimes include embezzlement, antitrust violations, and counterfeiting. See infra Sections I.B (explaining how the FBI categorizes white-collar crime), II.A (explaining how the U.S. Sentencing Commission categorizes white-collar crime), and II.B, Table 1 (showing that fraud makes up almost eighty percent of cases in the data). For centuries and up to present day, Anglo-American legal systems have tolerated frauds committed by the rich and powerful while systematically prosecuting poor and middle-class people for fraud offenses.32See, e.g., Emily Kadens, The Persistent Limits of Fraud Prevention in Historical Perspective, 118 Nw. U. L. Rev. 167, 173-79 (2023) (describing challenges in efforts during the Middle Ages to regulate fraud in consumer markets). But wealthy people have always committed fraud and other financial crimes even if they went unpunished. For example, the term “robber barons” originated to describe medieval English nobles who engaged in extortion.33Barbara A. Hanawalt, Fur-Collar Crime: The Pattern of Crime Among the Fourteenth-Century English Nobility, 8 J. Soc. Hist. 1, 1 (1975). The title of Hanawalt’s article refers to legislation by King Edward III of England that only permitted noble families to wear minever fur. See id. at 2. As historian Barbara Hanawalt describes, “kings and barons [of medieval England] both assumed that a certain amount of criminal activity was involved in being a noble and that it would be tolerated as long as it did not become excessive.”34Id. at 2; see id. at 3, 15 n.9 (reporting that 14 out of around 10,500 felony indictments in the fourteenth century involved members of the nobility). Although medieval English nobles engaged in “widespread extortion,” they were rarely criminally prosecuted.35Id. at 2–3 (noting that “the kings could use a number of informal and indirect means to control the illegal activities of their barons without bringing them into common criminal courts”); see also Kadens, supra note 32, at 168 (“Fraud is not, as it is sometimes assumed, a creature of modern capitalism, industrialization, the spread of complex financial systems, or the development of the corporation. On the contrary, many of the same types of frauds that we see today have existed throughout the history of organized society.”).

In other words, society saw financial crimes committed by the elite as part of the social fabric. Fraud was thus considered what observers would come to call a “street crime,” meaning it was viewed as a crime when committed by poor or middle-class people. For example, one of early America’s most infamous fraudsters—Charles Ponzi—was a poor immigrant from Italy who worked as a dishwasher, waiter, and bank teller before launching the eponymous scheme that would eventually result in his arrest, conviction of federal mail fraud, and a seven-year prison sentence.36Sewell Chan, A Look Back at Charles Ponzi the Schemer, N.Y. Times (Dec. 15, 2008, 12:53 PM), https://archive.nytimes.com/cityroom.blogs.nytimes.com/2008/12/15/ponzi-the-schemer-evoked-once-again [https://perma.cc/L842-PRXA]. Despite eventually amassing enormous wealth through his pyramid scheme, Ponzi was never a member of the elite.37Id. (quoting Mitchell Zuckoff describing, “[Ponzi] had his nose pressed against the glass . . . . He was not linked with Wall Street and New York, though he had dreams of being like Rockefeller”).

Meanwhile, as centuries went on, the term “robber barons” adapted to refer to business magnates of the nineteenth century who monopolized industries, corrupted government, engaged in unethical business practices, and exploited workers and investors.38See Hal Bridges, The Robber Baron Concept in American History, 32 Bus. Hist. Rev. 1, 1 (1958). Like the medieval robber barons whose criminal activity was ignored by the King,39See supra note 33 and accompanying text. the robber barons of the 1800s were also rarely prosecuted.40Lawrence M. Friedman, Crime and Punishment in American History 290 (1993) (“[T]here was a certain lack of zeal for punishing business behavior [before the 1930s].”) (cited in Eisinger, supra note 6 at 59).

By the early twentieth century and spurred by the Great Depression, the public and federal government grew increasingly interested in regulating markets and prosecuting members of the upper classes. During this era, Congress passed antitrust laws and laws regulating Wall Street.41Congress passed the Sherman Act in 1890, the Federal Trade Commission Act (creating the FTC) in 1914, and the Clayton Act in 1914. The Antitrust Laws, Fed. Trade Comm’n, https://www.ftc.gov/advice-guidance/competition-guidance/guide-antitrust-laws/antitrust-laws [https://perma.cc/JZ6J-S5TT]. As the Federal Trade Commission describes, “[w]ith some revisions, these are the three core federal antitrust laws still in effect today.” Id. Following the stock market crash of 1929, Congress in 1934 created the Securities and Exchange Commission (“SEC”) to restore confidence in the stock market and enforce securities laws.42Securities Exchange Act of 1934, Pub. L. No. 73-291, 48 Stat. 881 (creating the U.S. Securities and Exchange Commission and requiring stock exchanges to register with the federal government).

Scholars and the public needed an entirely new phrase—“white-collar crime”—to recognize that fraud committed by members of the elite was crime. Recognizing that members of the upper class engaged in enormous amounts of unpunished financial crime, sociologist Edwin Sutherland coined the term “white-collar crime” in his 1939 presidential address to the American Sociological Society.43Edwin H. Sutherland, White-Collar Criminality, 5 Am. Socio. Rev. 1, 1–2, n.1 (1940) (Thirty-Fourth Annual Presidential Address delivered at Philadelphia, Pa., Dec. 27, 1939). Sutherland went on to write a book by a similar name. Edwin H. Sutherland, White Collar Crime (1949).

Sutherland defined a “white-collar crime” as “a crime committed by a person of respectability and high social status in the course of his occupation.”44Sutherland, White Collar Crime, supra note 43, at 7. Sutherland’s basic thesis was that the academic methods by which crime was understood and measured at the time were invalid because “they have not included vast areas of criminal behavior of persons not in the lower class.”45Sutherland, White-Collar Criminality, supra note 43, at 2.

Sutherland critiqued the academic criminological community for focusing too heavily on “street crimes” perpetrated by “low status” people and for being insufficiently interested in crimes committed by people in “high status” occupations. As an example, Sutherland explained, “The ‘robber barons’ of the last half of the nineteenth century were white-collar criminals, as practically everyone now agrees.”46Id. Sutherland warned, however,

The present-day white-collar criminals . . . are more suave and deceptive than the “robber barons” . . . . Their criminality has been demonstrated again and again in the investigations of land offices, railways, insurance, munitions, banking, public utilities, stock exchanges, the oil industry, real estate, reorganization committees, receiverships, bankruptcies, and politics. Individual cases of such criminality are reported frequently, and in many periods more important crime news may be found on the financial pages of newspapers than on the front pages.47Id.

Beginning in the mid-twentieth century, the federal government began to articulate and attempt to carry out a new vision of white-collar prosecution. In the 1970s the SEC created its first enforcement division to uncover fraud.48Harwell Wells, The Securities and Exchange Commission’s Enforcement Division: A
History, Temple 10-Q, https://www2.law.temple.edu/10q/the-securities-and-exchange-commissions-enforcement-division-a-history [https://perma.cc/C8HF-X9MS].
In 1977, Congress passed the Foreign Corrupt Practices Act which outlawed bribery of foreign officials principally by large U.S. companies.49Foreign Corrupt Practices Act of 1977, Pub. L. No. 95-213, 91 Stat. 1494. In the 1980s, DOJ prosecuted over 1,000 cases associated with the savings and loan crisis, including some top executives at major banks.50Kitty Calavita, Henry N. Pontell & Robert H. Tillman, Big Money Crime: Fraud and Politics in the Savings and Loan Crisis 28 (1997) (“By the spring of 1992, in excess of one thousand defendants had been formally charged in major savings and loan cases, with a conviction rate of 91 percent . . . .”). During this time, as some observers noted, “Many U.S. Attorneys’ Offices . . . restructured their offices in order to develop and prosecute a large number of cases of white-collar crime.”51Kenneth Mann, Stanton Wheeler & Austin Sarat, Sentencing the White-Collar Offender, 17 Am. Crim. L. Rev. 479, 480 n.3 (1980); see also Elizabeth Hinton, From the War on Poverty to the War on Crime: The Making of Mass Incarceration in America 24 (2016) (noting that FBI crime data during the 1960s and 1970s “emphasized street crime to the exclusion of organized and white-collar crime”). The next subsection describes the mechanics of this modern era of federal enforcement of financial crime.

B.  Modern Fraud Prosecutions: 1990s Through Present

Efforts to differentiate financial crime committed by the elite from financial crime committed by poor or middle-class people were short-lived. Today, the term “white-collar” crime eludes easy definition.52Stuart P. Green, The Concept of White Collar Crime in Law and Legal Theory, 8 Buff. Crim. L. Rev. 1, 2 (2004) (claiming that “the meaning of white collar crime . . . is deeply contested. . . . [but d]espite its fundamental awkwardness, the term ‘white collar crime’ is now so deeply embedded within our legal, moral, and social science vocabularies that it could hardly be abandoned”). Scholars, journalists, and public officials often use the term as in its original definition—to refer to financial crimes committed by wealthy people in the course of business activity,53See infra note 111 and accompanying text. as exemplified by Ralph Nader’s pithy description of white-collar crime as “crime in the suites,” rather than “crime in the streets.”54Ralph Nader, White Collar Fraud; America’s Crime Without Criminals, N.Y Times, May 19, 1985 (§ 3), at 3, https://www.nytimes.com/1985/05/19/business/white-collar-fraud-america-s-crime-without-criminals.html [https://perma.cc/E7DE-SZQS].

However, official definitions of the term “white-collar” crime typically do not refer to the social status or occupation of those who perpetrate it, but rather, to the type of criminal behavior committed by the defendant.55The FBI explains that it would be impractical for the FBI to report white-collar crime statistics based on the offender’s socioeconomic status because that data is not available in the Uniform Crime Reports. See Cynthia Barnett, U.S. Dep’t of Just., Fed. Bureau of Investigation, The Measurement of White-Collar Crime Using Uniform Crime Reporting (UCR) Data 1 (2000) (“Although it is acceptable to use socioeconomic characteristics of the offender to define white-collar crime, it is impossible to measure white-collar crime with UCR data if the working definition revolves around the type of offender. There are no socioeconomic or occupational indicators of the offender in the data.”). The FBI, for example, defines “white-collar crime” as “those illegal acts which are characterized by deceit, concealment, or violation of trust and which are not dependent upon the application or threat of physical force or violence.”56Id. The National Incident-Based Reporting System (“NIBRS”), which compiles data on crimes reported to law enforcement, classifies the following crimes as white-collar crimes: fraud, bribery, counterfeiting/forgery, embezzlement, and writing bad checks.57Id. at 2.

This Article roughly follows the NIBRS definition but uses the term financial crime because, as this Article shows, the term white-collar crime is a misnomer. I define a crime as a financial crime if it is categorized as an antitrust violation, bribery, counterfeiting, forgery, fraud, or tax offense.58See infra Section II.A (describing how the data is constructed). Since the mid-1990s, the federal government has prosecuted around 10,000 financial crimes per year, most of them frauds.59See infra Appendix Figure A.1. The statistics presented in the Article show the same patterns when the data is restricted to fraud cases. Until fiscal year 2018, the U.S. Sentencing Commission reported separately whether a defendant’s offense of conviction was a fraud, larceny, or embezzlement. Beginning in 2018, however, the U.S. Sentencing Commission began combining these three types of crime into one category in the data. To make the data consistent throughout, I combined the three categories together under the label “financial crime” in the years prior to 2018. This section describes in broad terms how the federal government prosecutes and talks about financial crime.

1.  The Statutory Landscape

Federal law today defines many types of financial crimes, most of which are contained in Chapter 47 of Title 18 of the United States Code. The most commonly prosecuted federal financial crimes are embezzlement of public money, mail and wire fraud, bank fraud, and tax fraud.60See infra Appendix Table A.1. Congress has repeatedly expanded the scope of federal financial criminal law and, over the years 1994 to 2019, federal defendants were prosecuted for violations of many different types of fraud.61See id.

Federal prosecutors use mail fraud (and its sister crime, wire fraud) particularly expansively. The original mail fraud statute prohibited the use of the mails to advance “any scheme or artifice to defraud.”62Act of June 8, 1872, Pub. L. No. 42-335, § 301, 17 Stat. 283, 323 (revising, consolidating, and amending the statutes relating to the Post Office Department). Congress has expanded the mail fraud statute several times since its original passage. Mail fraud is now defined in 18 U.S.C. § 1341. The original purpose of the statute was to protect the U.S. Postal Service from being used to commit fraud. Mail was the “first communications network in the United States,”63Anuj C. Desai, Wiretapping Before the Wires: The Post Office and the Birth of Communications Privacy, 60 Stan. L. Rev. 553, 553 (2007). and in 1870 the U.S. Postal Service enjoyed a natural monopoly over mail delivery.64See id. at 573. Perhaps because the mail was so widely used, “[o]ver time, the mail fraud statute came to be viewed as a stop-gap provision that provides a ‘first line of defense’ to combat innovative frauds until Congress could enact more specific legislation.”65Peter J. Henning, Maybe It Should Just Be Called Federal Fraud: The Changing Nature of the Mail Fraud Statute, 36 B.C. L. Rev. 435, 437 (1995).

In 1995, Peter Henning contended that “the mail fraud statute has become the primary provision to extend federal jurisdiction to crimes traditionally prosecuted only at the state and local level.”66Id. Today nearly all frauds use mail, telephone, radio, or the Internet in some way, giving the federal government the ability to prosecute almost any fraud it chooses. Federal prosecutors exercise enormous discretion in deciding which fraud crimes to prosecute, and the resulting prosecutions therefore reflect decisions by prosecutors and law enforcement agents about which cases to prioritize.

Although there are many federal financial crimes, their defining characteristic is that they involve dishonesty. To this end, most financial crimes include mens rea elements that require the government to specifically prove the defendant’s deceitful intent.67Some observers point out that financial crime’s traditionally high mental state requirements have, to some extent, been eroded with theories of, for example, willful blindness or reckless regard for falsity. Baer, supra note 6, at 30-31 (2023). For example, the mail fraud statute requires proof that the defendant devised or intended a “scheme or artifice to defraud.”6818 U.S.C. § 1341. Health care fraud similarly requires proof that the defendant knowingly and willfully executed “a scheme or artifice . . . to defraud any health care benefit program” or to obtain, “by means of false or fraudulent pretenses, . . . any of the money or property owned by, or under the custody or control of, any health care benefit program.”69Id. § 1347 (a)(1)–(2).

Despite this common element, the financial crimes that are prosecuted vary widely on many grounds. Victims of financial crimes can be individuals, organizations, or the government. Some financial crimes have a single concrete victim, others have many, and yet others have no concrete victim (like insider trading). Some financial crimes involve wrongdoing that is also investigated and enforced by the government through civil proceedings (such as securities fraud or tax fraud), while others have no regulatory counterpart (such as embezzlement). The next section broadly describes how federal prosecutors and agents investigate and bring financial crime cases.

2.  Federal Prosecutions in Practice

Nearly all federal financial crime prosecutions are brought by prosecutors who work in the ninety-three U.S. Attorney’s Offices (“USAOs”). Each USAO is associated with exactly one of the 94 geographically distinct federal district courts, with one exception.70The District of Guam and the District of the Northern Mariana Islands share a USAO. Every USAO is led by a U.S. Attorney, who is appointed by the President. The prosecutors who work in USAOs are called Assistant United States Attorneys (“AUSAs”).

Although USAOs must follow centralized policies dictated by DOJ leadership, they for the most part work independently, prosecuting crimes that occur within their jurisdictions. Most prosecutorial decisions (such as the decision to bring criminal charges) are subject to little judicial oversight and courts are “hesitant to examine the decision whether to prosecute.”71Wayte v. United States, 470 U.S. 598, 608 (1985). As a result, prosecutors enjoy broad discretion in deciding how to carry out their work.72See Stephanos Bibas, Prosecutorial Regulations Versus Prosecutorial Accountability, 157 U. Pa. L. Rev. 959, 959 (2009) (“Few regulations bind or even guide prosecutorial discretion, and fewer still work well.”); William J. Stuntz, The Pathological Politics of Criminal Law, 100 Mich. L. Rev. 505, 506 (2001) (describing prosecutors as “the criminal justice system’s real lawmakers”). In theory, a defendant can challenge their prosecution on the ground that it was brought selectively—that is, based on a prohibited consideration such as the defendant’s race or religion. See Oyler v. Boles, 368 U.S. 448, 456 (1962). In practice, however, selective prosecution challenges virtually never succeed. See Richard H. McAdams, Race and Selective Prosecution: Discovering the Pitfalls of Armstrong, 73 Chi.-Kent L. Rev. 605, 615–16 (1998) (noting that since 1886 there has been only one published case dismissing a criminal charge based on racially selective prosecution). But see Alison Siegler & William Admussen, Discovering Racial Discrimination by the Police, 115 Nw. U. L. Rev. 987, 987 (2021) (describing how federal courts can and should lower the discovery standards for defendants alleging racial discrimination by the police).

Despite limited oversight from the courts, individual prosecutors are subject to other forms of workplace oversight. AUSAs are governed by the Justice Manual, which contains detailed rules for how individual prosecutors should exercise their discretion. For example, the Manual dictates that charging decisions should be reviewed by supervisors and specifies that “[a]ll but the most routine indictments should be accompanied by a prosecution memorandum that identifies the charging options supported by the evidence and the law and explains the charging decision[s] therein.”73U.S. Dep’t of Just., Just. Manual § 9-27.300 (2023).

The Manual also expresses a nationwide policy that federal prosecutors should usually charge “the most serious offense that is encompassed by the defendant’s conduct and that is likely to result in a sustainable conviction.”74Id. However, the Manual leaves room for an AUSA to deviate from this policy by also considering “whether the consequences of those charges for sentencing would yield a result that is proportional to the seriousness of the defendant’s conduct, and whether the charge achieves [the] purposes of the criminal law.”75Id.

Given these policies, how do prosecutors decide which cases to charge? The answer is complicated and varied, but much legal and sociolegal scholarship has shown the perhaps unremarkable phenomenon that prosecutors seem to like to bring cases they think they can win.76See Brandon Hasbrouck, The Just Prosecutor, 99 Wash. & Lee U. L. Rev. 627, 632 (2021) (“The adversary system derails many prosecutors, including progressive prosecutors, and turns them into win-seekers instead of neutral agents of justice.”); Rachel E. Barkow, Institutional Design and the Policing of Prosecutors: Lessons from Administrative Law, 61 Stan. L. Rev. 869, 883 (2009) (suggesting that prosecutors “may feel the need to be able to point to a record of convictions and long sentences if they want to be promoted or to land high-powered jobs outside the government” and prefer to “keep up [their] conviction rate”); Tracey L. Meares, Rewards for Good Behavior: Influencing Prosecutorial Discretion and Conduct with Financial Incentives, 64 Fordham L. Rev. 851, 867 (1995) (“A prosecutor will naturally select the stronger cases to charge.”). But see Richard T. Boylan, What Do Prosecutors Maximize? Evidence from the Careers of U.S. Attorneys, 7 Am. L. & Econ. Rev 379, 379 (2005) (finding that “conviction rates do not appear to affect the careers of U.S. attorneys”). This is because obtaining convictions is often a metric for promotion and advancement.77Stephanos Bibas, Plea Bargaining Outside the Shadow of Trial, 117 Harv. L. Rev. 2463, 2471 (2004) (“[P]rosecutors want to ensure convictions . . . . Favorable win-loss statistics boost prosecutors’ egos, their esteem, their praise by colleagues, and their prospects for promotion and career advancement.”). Winning cases is also important for appropriations. As Lauren Ouziel describes,

U.S. Attorney’s Offices, after all, need money, and federal funds are not forthcoming—either from Congress in the first instance or Main Justice in the subsequent allocation—without some measure of demonstrated performance. For federal prosecutors, the relevant performance metrics are defendants charged and convicted. Both of these metrics determine the lump sum congressional appropriation for all ninety-three U.S. Attorneys’ Offices across the country, while individual offices’ caseloads largely determine the allocation of those funds among the offices. In short, case volume and prosecutorial success dictate a U.S. Attorney’s Office’s budget allocation.78Lauren M. Ouziel, Ambition and Fruition in Federal Criminal Law: A Case Study, 103 Va. L. Rev. 1077, 1108–09 (2017) (citing U.S. Dep’t of Justice, U.S. Attorneys, FY 2014 Performance Budget Congressional Submission 1, 15; Dep’t of Justice, Office of Inspector Gen., Audit Div., Audit Report 09-03, Resource Management of United States Attorneys’ Offices 7–10 (Nov. 2008)).

After a person is convicted of a federal crime, federal judges sentence them. At sentencing, a judge can impose fines or imprisonment or both on a defendant, and some scholars have pointed out that fines are imposed more frequently in financial crime prosecutions than in other federal prosecutions.79Max Schanzenbach & Michael L. Yaeger, Prison Time, Fines, and Federal White-Collar Criminals: The Anatomy of a Racial Disparity, 96 J. Crim. L. & Criminology 757, 768 (2006). Some have theorized that fines are more appropriate for defendants convicted of financial crimes because their crimes are more deterrable.80See, e.g., Stephanos Bibas, White-Collar Plea Bargaining and Sentencing After Booker, 47 Wm. & Mary L. Rev. 721, 724 (2005) (“An economist would argue that if one increased the expected cost of white-collar crime by raising the expected penalty, white-collar crime would be unprofitable and would thus cease.”). Others, including Richard Posner, have argued that fines should be more widely used for the entire spectrum of crimes given the high costs of physical incarceration. See Richard A. Posner, Optimal Sentences for White-Collar Criminals, 17 Am. Crim. L. Rev. 409, 409–10 (1980) (arguing in favor of “the substitution, whenever possible, of the fine (or civil penalty) for the prison sentence as the punishment for crime”). But see Dorothy S. Lund & Natasha Sarin, Corporate Crime and Punishment: An Empirical Study, 100 Tex. L. Rev. 285, 285 (2021) (arguing that “enforcers are unlikely to achieve optimal deterrence using fines alone”); Jed S. Rakoff, The Financial Crisis: Why Have No High-Level Executives Been Prosecuted?, N.Y. Rev. Books (Jan. 9, 2014), https://www.nybooks.com/articles/2014/01/09/financial-crisis-why-no-executive-prosecutions [https://perma.cc/5BRX-UBAD] (arguing that fines are inadequate to change corporate behavior and that the threat of imprisonment against executives would be a more effective deterrent). Researchers have also argued that the prevalence of fines in financial crime sentencing reflects a fine/incarceration tradeoff, in which the greater a defendant’s ability to pay a fine, the less (if any) imprisonment is imposed at sentencing.81See Joel Waldfogel, Are Fines and Prison Terms Used Efficiently? Evidence on Federal Fraud Offenders, 39 J.L. & Econ. 107, 107 (1995).

The literature on white-collar crime’s fine/incarceration tradeoff might give the impression that fines are widespread in financial crime prosecutions, but this is not the case. Most federal financial crime defendants do not have any fines imposed in their cases. In the data, the median fine amount for a defendant convicted of a federal financial crime is $0.82See infra Table 1. A fine of just $500 represents the top thirteen percent of fines imposed among people convicted of financial crimes.83This observation is based on the author’s analysis of the data. See Didwania, Data, supra note 17. It is true that fines are more prevalent among financial crime defendants than others (a fine of $500 for a federal defendant convicted of a non-financial crime would represent the top nine percent of all fines imposed),84Id. but it is not the case that fines are widespread among those who are convicted of financial crimes. Instead, fines are much more relevant in cases involving corporate defendants. This is because corporations cannot be imprisoned, fines generate revenue, and prosecutors worry about the collateral consequences that criminal conviction can impose on large corporations.85For example, Mary Jo White, former U.S. Attorney for the Southern District of New York (and future SEC Chair) said in an interview, “[a]ny prosecutor hesitates before bringing an action against a company because of the fear that that company will go out of business.” Interview with Mary Jo White, Debevoise, New York, New York, 19 Corp. Crime Rep. (Dec. 12, 2005), https://www.corporatecrimereporter.com/news/200/category/sampleinterviews [https://perma.cc/MLL3-UCCM].

In addition to fines and imprisonment, convicted defendants will usually be ordered to pay restitution to any concrete victim. Restitution is different from a fine. A fine is a form of punishment imposed on a defendant and usually paid to the government prosecuting the case. Restitution is instead paid by the defendant to either the victim or a government restitution fund. Like the law in all states, federal law requires courts to order restitution in any case “in which an identifiable victim or victims has suffered a physical injury or pecuniary loss.”8618 U.S.C. § 3663A(c)(1)(B).

Unlike fines, most defendants convicted of a financial crime are ordered to pay some restitution. The median restitution amount ordered is around $6,000.87See infra Table 1. In contrast, for federal defendants convicted of non-financial crimes, fewer than ten percent are ordered to pay any restitution.88This observation is based on the author’s analysis of the data. See Didwania, Data, supra note 17.

The majority of defendants convicted of federal financial crimes are sentenced to prison. Sentences for financial crime defendants are lower than the average among other types of federal crimes. For federal criminal defendants convicted of financial crimes, the average sentence is around sixteen months.89See infra Table 1. For all other federal criminal defendants, the average sentence is fifty-three months.90This observation is based on the author’s analysis of the data. See Didwania, Data, supra note 17. This could reflect the fact that most financial crimes do not carry mandatory minimum penalty provisions.91The only type of financial crime that carries a mandatory minimum is identity theft. Aggravated identity theft includes a two-year mandatory minimum penalty. 18 U.S.C. § 1028A; see also An Overview of Mandatory Minimum Penalties in the Fed. Crim. Just. Sys. § 3 (U.S. Sent’g Comm’n 2017) (listing federal crimes that carry mandatory minimum penalties).

3.  Federal Financial Crime Archetypes

This subsection illustrates some of the kinds of financial crime cases the federal government prosecutes. It centers around four real-world examples of federal financial crimes, from least to most severe.92As Miriam Baer has pointed out, most federal fraud offenses are not statutorily graded the way other types of crimes are. See Miriam H. Baer, Sorting Out White-Collar Crime, 97 Tex. L. Rev. 225, 228 (2018). Instead, a federal fraud’s severity is largely driven by the dollar amount of loss, as dictated by § 2B1.1 of the United States Sentencing Guidelines. See id. at 250 (“Because the federal criminal code declines to differentiate fraud up front—either by amount, mens rea, or degree of risk—whatever sorting there is of fraud offenses takes place at sentencing.”). These cases exemplify nationwide patterns that this Article reports and explores in Part II, and this Article returns to these examples throughout.

In Case A, a man who is a citizen of Mexico used a social security number belonging to another person to secure employment and attend a job orientation training with a local company.93Press Release, U.S. Attorney’s Office for the Eastern District of Louisiana, Mexican National Sentenced for Illegally Using a Social Security Number Belonging to Another Person (Oct. 12, 2022), https://www.justice.gov/usao-edla/pr/mexican-national-sentenced-illegally-using-social-security-number-belonging-another [https://perma.cc/4DMK-8C7S]. The man was prosecuted in the Eastern District of Louisiana and was ultimately convicted of violating 18 U.S.C. § 408(a)(7)(b), which makes it a crime to fraudulently use another person’s social security number. The man was sentenced to one year of probation.

In Case B, a man received Social Security and Department of Defense benefits intended for his late father for four years after his father’s death.94Press Release, U.S. Attorney’s Office for the Southern District of Ohio, Fifteenth Person Charged with Theft in Ongoing Social Security Benefits Fraud Investigation (Aug. 10, 2020), https://oig.ssa.gov/news-releases/2020-08-10-audits-and-investigations-investigations-aug4-oh-fifteenth-person-charged-social-security-fraud [https://perma.cc/5CFK-7JWP]. The man’s elderly father had moved in with the man in 2012.95Sentencing Memorandum of Defendant Napoleon Crawford at 2, United States v. Crawford, No. 1:20CR029 (S.D. Ohio Aug. 6, 2021). The man cared for his father for next four years, until his father’s death at age 92 in 2016.96Id. When the man began caring for his father in 2012, they joined bank accounts, into which his father’s benefits were deposited.97Id. After his father’s death, a death certificate was properly filed, but his late father’s benefit payments continued to be deposited into their joint bank account.98Id. Over the four years that followed his father’s death, the man collected $42,103 in Social Security benefits and $41,609 in Department of Defense benefits to which he was not entitled.99Press Release, supra note 94.

Case B was prosecuted in the U.S. District Court for the Southern District of Ohio. The man pled guilty to theft of public money. He was sentenced to eight months in prison and ordered to pay $83,712 in restitution to the Social Security Administration (“SSA”) and Department of Defense.100Id.

Case B was part of a federal initiative called the Social Security Administration Fraud Prosecution Project.101Id. The SSA Fraud Prosecution Project is a collaboration of the SSA Office of the Inspector General (“SSA OIG”) and DOJ.102Id. The investigation of Case B also involved employees of the Department of Defense Office of Inspector General, the Veteran’s Administration Office of Inspector General, the United States Office of Personnel Management Office of Inspector General, and the United States Secret Service.103Id. It appears that many federal agencies and employees devoted significant resources to bringing Case B and others like it.

In all, the SSA OIG reports that as a result of its audit program, it discovered dozens of instances of people collecting social security or veteran benefits intended for another person in Ohio, a state that has an adult population of more than eight million.104Id. See also Gustafson, supra note 14, at 57 (finding that in California, the state conducts biometric imaging (that is, fingerprinting) of all welfare applicants as a way to detect fraud and discovers around three people per month who have submitted a duplicate application). The SSA OIG investigation has led the USAO for the Southern District of Ohio to prosecute at least fifteen people in cases like Case B. The losses to SSA associated with these cases average just under $60,000 per defendant.105Id.

In Case C, a married couple owned and operated a company called Kingdom Connected Investments (“KCI”), which they advertised as a Christian organization.106Press Release, U.S. Attorney’s Office for the District of South Carolina, Married Greenville Business Owners Sentenced to More than Seventeen Total Years, Ordered to Pay More than $2.5 Million in Restitution for Defrauding Home Buyers and Sellers (Oct. 5, 2020), https://www.justice.gov/usao-sc/pr/married-greenville-business-owners-sentenced-more-seventeen-total-years-ordered-pay-more [https://perma.cc/SAN6-ZCTE]. KCI sought to pair clients who fell into two categories: (1) homeowners who owed more on their homes than the home was worth (that is, they were “underwater” on the home); and (2) potential homebuyers who did not have a high enough credit score to qualify for a conventional mortgage. KCI operated by matching homeowners (sellers) and buyers. KCI told the sellers they would transfer title of the home to KCI and take over the home’s mortgage payments, allowing the homeowners to get out of their underwater mortgage. KCI collected down payments from the buyers, telling them they were renting-to-own the home.

None of this was true. In reality, KCI never actually purchased the sellers’ homes, which meant each property still had an existing mortgage in the seller’s name(s) after the sellers thought they no longer owned the home. Rather than using the buyers’ down payments to pay the mortgages in full as promised, KCI used much of these down payments for personal use and to try to build their real estate business. Eventually, with the mortgages unpaid, nearly all the homes went into foreclosure and sold at auction. Many sellers learned that KCI had not actually purchased their home when they received foreclosure notices. Many of KCI’s buyers, who thought they were renting-to-own their homes, learned the truth when the home’s new owners sought to evict them. In all, KCI received $2.7 million from the buyers but only made $1.4 million in mortgage payments. Approximately 130 properties were involved in the scam, suggesting the average buyer lost around $20,000. Most sellers had their credit scores ruined by the foreclosures.

Case C was prosecuted in the U.S. District Court for the District of South Carolina. A federal jury found the defendants guilty of conspiracy to commit mail fraud and equity skimming after just ninety minutes of deliberation. The husband and wife were sentenced to seventy-eight and 136 months in prison, respectively, and ordered to pay $2,664,796.69 in restitution.

Case D will be familiar to many readers. JPMorgan Chase, a major U.S. bank, knowingly packaged shoddy mortgages into securities that did not meet its credit standards. JPMorgan Chase sold these securities to investors. A JPMorgan Chase manager (and attorney), Alayne Fleischmann, described JPMorgan Chase’s mortgage securities business as a “massive criminal securities fraud.”107Matt Taibbi, The $9 Billion Witness: Meet JPMorgan Chase’s Worst Nightmare, Rolling Stone (Nov. 6, 2014), https://www.rollingstone.com/politics/politics-news/the-9-billion-witness-meet-jpmorgan-chases-worst-nightmare-242414 [https://perma.cc/SWB2-6ARH?type=standard]. Before the 2008 crash, Fleischmann wrote a thirteen-page memo to her supervisor warning that the bank was improperly packaging bad mortgages into securities and selling them as investments. Fleischmann was fired and bankers at JPMorgan Chase continued in their scheme. Fleischmann eventually became a whistle-blower and provided detailed evidence about JPMorgan Chase’s wrongdoing to the SEC and federal prosecutors.

Unlike the defendants in Cases A, B, and C, the federal government never prosecuted either JPMorgan Chase the organization or any of its employees for their fraud. Chase instead agreed to a $13 billion settlement with federal and state agencies for wrongdoing during the crisis. As a publicly traded company, Chase paid the settlement with shareholders’ money and the settlement agreement did not name any bankers. A few weeks later, Chase’s CEO, Jamie Dimon, received a seventy-four percent raise, bringing his salary to $20 million per year.

C.  How We Talk About Financial Crime

Academic and journalistic writing about white collar crime tends to focus on cases like D.108See supra text accompanying note 7. It examines and seeks to understand the causes and consequences of a criminal system that is unwilling or unable to convict large firms and the people who lead them, even when those firms and people create staggering social harm and there is evidence that their conduct violates the criminal law. Much work in this area documents the DOJ’s increased use of deferred and non-prosecution agreements for companies engaged in corporate crime.109See, e.g., Arlen & Kahan, supra note 7; Veronica Root Martinez, The Government’s Prioritization of Information Over Sanction: Implications for Compliance, 83 L. & Contemp. Probs. 85, 85–87 (2020). Other work asks similar questions about individuals who hold positions of leadership in corporate organizations that commit crimes.110In this vein, some recent scholarship about white-collar crime committed by individuals has focused on a 2015 Memo from Deputy Attorney General Sally Quillian Yates (the “Yates Memo”) that outlines steps that federal prosecutors should take to “strengthen [the] pursuit of individual corporate wrongdoing.” Memorandum from Deputy Att’y Gen. Sally Quillian Yates to Assistant Att’ys Gen. & All U.S. Att’ys., Individual Accountability for Corporate Wrongdoing (Sept. 9, 2015) (on file with DOJ). For example, some have pointed out that even after the Yates Memo was promulgated, DOJ continued to enter deferred prosecution agreements with corporations without charging individuals. See, e.g., Paola C. Henry, Individual Accountability for Corporate Crimes After the Yates Memo: Deferred Prosecution Agreements & Criminal Justice Reform, 6 Am. U. Bus. L. Rev. 153, 160–161 (2016) (describing the post-Yates Memo case in which General Motors employees intentionally failed to disclose a safety defect in their ignition switches, which led to at least 124 deaths, but federal prosecutors entered a deferred prosecution agreement with GM without charging any individuals).

In contrast to much of the literature, this Article focuses instead on cases like A, B, and C, which represent the bread and butter of most federal financial criminal enforcement in the United States. Many scholarly examinations of federal white-collar crime characterize these cases as not white-collar crime. For example, Samuel Buell explains in his 2014 study of white-collar sentencing:

Many white collar offenses, maybe even most of them, are committed by pedestrian hucksters, scam artists, cheaters, and liars. Such persons have been among us for ages. This Article makes few claims about the treatment of this class of offenders—the home buyer who lies to obtain a mortgage, the taxpayer who cheats the Internal Revenue Service (IRS), the restaurant manager who bribes the health inspector, and their ilk. The discussion here responds to a public debate that does not often mention the small-time crook.111Buell, supra note 7, at 830–31 (2014); see also Mihailis E. Diamantis, White-Collar Showdown, 102 Iowa L. Rev. 320, 320 (2017) (“Not many people would rank white-collar criminals among the downtrodden of the criminal justice system.”); Darryl K. Brown, Street Crime, Corporate Crime, and the Contingency of Criminal Liability, 149 U. Pa. L. Rev. 1295, 1315 (2001) (“Painting with an overbroad brush, street offenders are outside the mainstream norms of society. More committed to subcultures or simply irrational, violent, or greedy, their crimes are clearly intentional. White-collar offenders, on the other hand, except for those white-collar crimes that plainly mimic street crimes—for example, embezzling from an employer is stealing and credit card or insurance fraud are just other forms of theft—are more reasonable, mainstream people.”). But see Pedro Gerson, Less is More?: Accountability for White-Collar Offenses Through an Abolitionist Framework, 2 Stet. Bus. L. Rev. 144, (noting that “[a]n important caveat to note at the outset is that [the author’s] definition of white-collar crime is significantly narrower than the one used by law enforcement, which focuses on the type of offenses and centers on crimes of ‘deceit, concealment or violation of trust’ without the use of force”); Benjamin Levin, Wage Theft Criminalization, 54 U.C. Davis L. Rev. 1429, 1483-84 (2021) (noting that the sorts of incidents reported in a 2000 FBI report tended to be low-level property crimes and frauds rather than “the dominant cultural (and legal) imagination of ‘white-collar crime’ ”); Daniel Richman, Federal White Collar Sentencing in the United States: A Work in Progress, 76 L. & Contemp. Probs. 53, 53 (2013) (“[C]rimes involving fraud, deceit, theft, embezzlement, insider trading, and other forms of deception . . . include[] a great many offenders and offenses of the middling sort.”); Posner, supra note 80, at 409–10 (using the term white-collar crime “to refer to the nonviolent crimes typically committed by either (1) well-to-do individuals or (2) associations, such as business corporations and labor unions, which are generally ‘well-to-do’ compared to the common criminal”).

This Article argues that when—as Buell notes—the public debate about white-collar crime excludes financial crimes committed by people who are not wealthy executives, the exclusion is not merely semantic. Using the term “white-collar crime” to only include prosecutions of elite people shields from public view the vast majority of prosecutions that happen under our financial criminal laws.

We have not always talked about financial crime this way. This Article provides updated and more comprehensive answers to some of the questions asked in a series of studies produced in the 1980s through early 2000s by Stanton Wheeler and others called the Yale Studies on White-Collar Crime (“Yale Studies”). In the final of four studies in this series, the authors analyzed the personal characteristics of those whom the authors characterized as federal white-collar defendants. Using a sample of roughly 210 white-collar defendants randomly sampled from seven federal district courts, the authors found that their sample of white-collar defendants “departs from common images of the typical white collar offender in that they are very similar to average or middle class Americans.”112David Weisburd, Elin Waring & Ellen Chayet, U.S. Dep’t of Just., White Collar Crime and Criminal Careers 2 (1993) (citing David Weisburd, Stanton Wheeler, Elin Waring & Nancy Bode, Crimes of the Middle Classes: White Collar Offenders in the Federal Courts (1991)). The seven districts studied were: the Central District of California, the Northern District of Georgia, the Northern District of Illinois, the District of Maryland, the Southern District of New York, the Northern District of Texas, and the Western District of Washington. Id. The authors also noted that their study found white-collar crimes to “have a much more mundane quality than those which are associated with white collar crime in the popular press,” noting that “the bulk of white collar crimes prosecuted in the federal courts are undramatic and maybe committed by people of relatively modest social status.”113Id. at 11.

The Yale study’s findings are similar but less extreme than the updated and more fulsome patterns this Article documents in Part II. This Article, for example, suggests that the average financial crime defendant is likely to have lower income than the average U.S. adult, whereas the authors of the Yale study find that “most white-collar offenders were from the middle class, that is, they were significantly above the poverty line, but they were not from the upper echelons of wealth and social status.”114David Weisburd, Stanton Wheeler, Elin Waring & Nancy Bode, Crimes of the Middle Classes: White Collar Offenders in the Federal Courts, U.S. Dep’t of Just., Off. of Just. Programs (1991), https://ojp.gov/ncjrs/virtual-library/abstracts/crimes-middle-classes-white-collar-offenders-federal-courts [https://perma.cc/VP9D-W3H4]. Part II also shows that Black people are disproportionately prosecuted for white-collar crimes, which the Yale study did not find.

A likely reason the nationwide findings presented in this Article suggest the federal financial criminal defendant population is even less advantaged than as suggested by the Yale study is that the Yale authors’ sample was not representative of all federal financial crime prosecutions. The authors explain that they chose seven districts “in part because some of them were known to have a significant amount of white-collar prosecution,”115Weisburd et al., supra note 112, at 16. and all of the chosen districts contain major U.S. cities. By focusing on districts with active and sophisticated white-collar dockets in large U.S. cities, the Yale study likely overrepresents the income of all federal financial crime defendants. It also uses a sample of federal financial crime defendants whose racial makeup (seventy-eight percent White) is different from what this Article observes in its nationwide analysis (forty-nine percent White).116Another possible explanation for this difference is that over time the federal government might have increasingly prosecuted low-income people for financial crimes. The Yale study considered defendants sentenced between 1976 and 1978; this Article considers defendants prosecuted in 1994 through 2019, so perhaps the federal government’s enforcement behavior changed in the sixteen years between our studies.

This Article also relates to Max Schanzenbach and Michael Yaeger’s 2006 examination of racial disparities in federal white-collar cases.117See Schanzenbach &Yaeger, supra note 79, at 758. Using regression analysis, Schanzenbach and Yaeger find that after controlling for many relevant defendant and case characteristics, Black and Hispanic defendants convicted of white-collar crimes receive longer prison sentences than do White defendants.118Id. at 790. They also find that a significant portion of this inequality can be explained by defendants’ ability to pay a fine, lending support to the idea that there is a fine/incarceration tradeoff in white-collar cases.119See id. at 792.

This Article fundamentally differs from Schanzenbach and Yaeger’s work because this Article is a descriptive analysis. Many studies—like Schanzenbach and Yaeger’s—estimate whether defendants within a criminal system appear to be treated differently for reasons they should not be (such as their race,120See, e.g., Crystal S. Yang, Free At Last? Judicial Discretion and Racial Disparities in Federal Sentencing, 44 J. Legal Stud. 75, 75 (2015). skin color,121See, e.g., Traci Burch, Skin Color and the Criminal Justice System: Beyond Black-White Disparities in Sentencing, 12 J. Empirical Legal Stud. 395, 395 (2015). gender,122See, e.g., Sonja B. Starr, Estimating Gender Disparities in Federal Criminal Cases, 17 Am. L. & Econ. Rev. 127, 127 (2015). or wealth123See, e.g., Christine S. Scott-Hayward & Henry F. Fradella, Punishing Poverty: How Bail and Pretrial Detention Fuel Inequalities in the Criminal Justice System 45 (2019).). In contrast, this Article does not seek to advance a causal claim about the sources of inequality. To that end, this Article does not compare the outcomes of federal financial crime defendants to each other; it compares the population of federal financial crime defendants to the underlying U.S. adult population. It then examines whether, where, and for how long these inequalities in who is prosecuted have existed. The next Part presents this empirical analysis.

II.  INEQUALITY IN FEDERAL FINANCIAL CRIME PROSECUTIOS

Between 1994 and 2019, 1.7 million defendants were convicted of federal crimes and sentenced under the U.S. Sentencing Guidelines.124This count does not reflect defendants who were convicted of offenses carrying a statutory maximum term of incarceration of six months or less (that is, petty misdemeanor cases), see U.S. Sent’g Guidelines Manual § 1B1.9 (U.S. Sent’g Comm’n 2021), which are typically handled by federal magistrate judges. 28 U.S.C. § 636(a)(4). Infra Part II. Around 15% of these defendants were convicted of financial crimes, making financial crime the third-most prosecuted type of federal crime over this period, following drug crime (35% of cases) and immigration crime (25% of cases).125This observation is based on the author’s analysis of the data. See Didwania, Data, supra note 17. Most defendants convicted of financial crimes were convicted of some type of fraud, and even counted alone, fraud is the third-most prosecuted type of federal offense.126Id.

This Part presents the first nationwide empirical analysis of federal financial crime cases. Section II.A explains how I constructed the data set. Section II.B presents summary information about federal financial crime cases. Sections II.C through II.E use sentencing data matched to county-level population data to examine inequality in who is prosecuted for federal financial crimes. Section II.C shows that people who are Black and low-income are overrepresented in financial crime prosecutions relative to the U.S. adult population, while people who are White and middle- to high-income are underrepresented. Section II.D shows that income and race gaps in the prosecution of financial crime have narrowed over the last few decades but remain significant. Section II.E documents differences in these inequality patterns across federal districts. It shows that USAOs in the Deep South prosecute female defendants at the highest rates. Because states in the Deep South have among the largest Black populations in the U.S., their more intensive prosecution of women for financial crimes drives the overrepresentation of Black women among financial crime defendants. Section II.E also shows that Black defendants are overrepresented in financial crime cases in nearly all federal districts, which demonstrates that the nationwide inequality patterns are not solely a function of different prosecutorial priorities between districts.

A.  Data

The descriptive analysis that follows presents two types of facts about federal financial crime prosecutions. First, it describes the scale of federal prosecution of financial crime. It answers questions like: How many people does the federal government prosecute for financial crimes per year? How does this number compare to prosecutions for other types of federal crimes? How has this number changed over time? Second, the analysis describes representation in federal prosecutions of financial crime. It answers questions like: Are low- or high-income people over- or underrepresented among federal defendants charged with financial crimes? Which, if any, racial or gender groups are over- or underrepresented? Does over- or under-representation vary over time? Does it vary between USAOs?

Answering these descriptive questions requires two types of data: data on federal criminal cases and data on the U.S. adult population. The dataset used in this Article includes quantitative data of the roughly 1.7 million federal defendants sentenced under the U.S. Sentencing Guidelines in fiscal years 1994 through 2019, matched at the district and year level to population data from the U.S. Census. I built the federal criminal case dataset by combining annual data files published by the U.S. Sentencing Commission (“Commission”).127The Commission data files are available for download from the U.S. Sentencing Commission website (fiscal years 2002–2021) and through the Inter-university Consortium for Political and Social Research (fiscal years 1987–2019). See Monitoring of Federal Criminal Sentences Series, Inter-university Consortium for Pol. and Soc. Rsch., https://www.icpsr.umich.edu/web/ICPSR/series/83 [https://perma.cc/8DN8-3EFT]; Commission Datafiles, U.S. Sent’g Comm’n, https://www.ussc.gov/research/datafiles/commission-datafiles [https://perma.cc/U2U7-NLYA]. To compute inequality statistics, I dropped from the dataset defendants whose race, Hispanic ethnicity, or gender information are reported as missing (roughly four percent of defendants).

The Commission data files include thousands of variables that describe federal criminal defendants and their cases. Critically for this project, the Commission data include a defendant’s self-reported race and Hispanic ethnicity, gender,128The Commission data uses a binary variable for gender (Male/Female), which the Codebook simply said “indicates the offender’s gender.” U.S. Sent’g Comm’n, Variable Codebook for Individual Offenders 31 (2013). For at least some of the 1994–2019 period, the Federal Bureau of Prisons’ Transgender Offender Manual indicated that an inmate’s gender identity, rather than their gender assigned at birth, be considered when recommending a housing facility, which suggests that transgendered prisoners are likely coded according to their gender identity rather than biological sex. See Daniel Politi, Trump Administration Gets Rid of Obama-Era Rules that Protected Transgender Inmates, Slate (May 13, 2018, 8:59 PM), https://slate.com/news-and-politics/2018/05/trump-administration-gets-rid-of-obama-era-rules-that-protected-transgender-inmates.html [https://perma.cc/PP8P-PPBH]. level of formal education, age, and the nature of the defendant’s prior criminal record. The Commission data also include variables that provide information about the subject of the defendant’s case, such as the type of offense (divided into thirty-five categories) and the statutes of conviction. The Commission data also include variables describing case outcomes, including details of the sentence imposed upon the defendant and their advisory sentencing range. Finally, the Commission data report the month, year, and federal district court in which the defendant was sentenced. These variables allow me to understand the geography and history of inequalities in federal financial crime prosecutions.

After building the Commission dataset, I merged it with county-level data published by the U.S. Census Bureau that describes the U.S. adult population (“Census Data”). The Census Data’s county-level intercensal population estimates include annual age-by-race-by-gender data of county populations.129See Annual County Resident Population Estimates by Age, Sex, Race, and Hispanic Origin: April 1, 2020 to July 1, 2019, (CC-EST2019-ALLDATA), U.S. Census Bureau, https://www.census.gov/data/tables/time-series/demo/popest/2010s-counties-detail.html [https://perma.cc/XEN9-83YG]; Intercensal Estimates of the Resident Population by Five-Year Age Groups, Sex, Race, and Hispanic Origin for Counties: April 1, 2000 to July 1, 2010, U.S. Census Bureau, https://www.census.gov/data/
datasets/time-series/demo/popest/intercensal-2000-2010-counties.html [https://perma.cc/XZ4R-FS93]; State and County Intercensal Datasets 1990–2000, U.S. Census Bureau, https://www.census.gov/data/datasets/time-series/demo/popest/intercensal-1990-2000-state-and-county-characteristics.html [https://perma.cc/9J3T-DNAA].
The Economic Research Service of the USDA publishes county-level educational attainment information of the adult population using data from the U.S. Census and American Community Survey.130See Educational Attainment for Adults Age 25 and Older for the U.S., States, and Counties, 1970–2020, USDA, Econ. Rsch. Serv., https://www.ers.usda.gov/data-products/county-level-data-sets/county-level-data-sets-download-data [https://perma.cc/7HBS-GZJ2]. Unlike population data, this data is not reported for every year. It is only reported for 1970, 1980, 1990, 2000, 2007–11 (five-year average), and 2016–2020 (five-year average). For this project, I use the data from 2007–2011 because it is closest to the midpoint of the study period (1994 to 2019).

After compiling the county-level Census Data, I aggregated it to the federal district level with a district‑to-county crosswalk file.131Mary Eschelbach Hansen, Jess Chen & Matthew Davis, United States District Court Boundary Shapefiles (1900–2000), Inter-univ. Consortium for Pol. & Soc. Res. (Mar. 2, 2015), https://doi.org/10.3886/E30468V1 [https://perma.cc/5NA7-94ZH]. This matched data allowed me to measure per capita prosecution rates between districts and to compare characteristics of the federal defendant population with the entire adult resident population over time and within each federal judicial district.

B.  Preliminary Descriptive Statistics of Federal Financial Crime Cases

Before examining inequality in federal financial crime cases, Table 1 presents descriptive statistics of these cases from the data. I define a case as a “financial crime” if the Commission data characterizes it as an antitrust, bribery, counterfeiting, forgery, fraud, embezzlement, larceny,132Although larceny is not typically considered a white-collar crime, I include it in my definition for consistency because in fiscal year 2018, the Commission data began combining fraud, embezzlement, and larceny into one offense category. Defendants coded as committing larceny crimes in years prior to 2018 were frequently convicted of fraud and embezzlement crimes. or tax crime.

The Commission data do not include a variable to characterize the victim(s) in the case, so I coded this variable based on the criminal statute under which the defendant was convicted. Based on the statute of conviction, I coded the case as involving one of these four victim types: (1) a government victim; (2) a private victim; (3) no concrete victim; or (4) an unknown victim. For example, a case in which the defendant is convicted of embezzling or stealing public money is coded as having a government victim.133See 18 U.S.C. § 641. A case in which the defendant is convicted of defrauding a bank is coded as having a private victim.134See id. § 1344. A case in which the defendant is convicted of making a false statement to a federal agent is coded as having no concrete victim.135See id. § 1001. A case in which a person is convicted of defrauding a health insurer is coded as having an unknown victim because a person can commit this crime by defrauding either a government insurer (like Medicare) or a private insurer.136See id. § 1347.  Appendix Table A.1 lists the statutory provisions for defendants convicted of the most common financial crimes and how they were coded.137A complete list of all statutory provisions and how they were coded is on file with the author and available by request.

Table 1 provides summary statistics of many variables about the defendants and their cases in the data. Column (1) of Table 1 presents averages for the variables across all 276,210 defendants convicted of financial crimes in the years 1994–2019. Columns (2) through (5) present averages for the same variables among defendants whose crimes involve the lowest losses (column (2)) through largest losses (column (5)).138The observations in columns (2) through (5) do not sum to 276,210 because the “loss amount” variable is only available beginning in 1999. Even beginning in 1999, around twenty percent of observations are missing an entry in this variable. Because the severity of financial crimes is (for the most part) increasing in loss amount, readers should think of moving across Table 1 from column (2) to column (5) as moving from less serious to more serious financial crimes.139It is important to note that when I use the term “loss,” throughout this Article, I mean the “dollar amount of loss for which the offender is held responsible,” which is how this variable is defined by the Commission. Commentary to the U.S. Sentencing Guidelines directs courts to consider “actual or intended loss,” and there appears to be a recent circuit split on the question of whether using intended loss is acceptable. Compare United States v. Gadson, 77 F.4th 16, 21–22 (1st Cir. 2023) (district court did not commit plain error by using intended loss to calculate bank-fraud defendant’s base offense level) with United States v. Banks, 55 F.4th 246, 248 (3d. Cir. 2022) (concluding that the Commission’s commentary that includes “intended loss” in the definition of “loss” should be afforded no weight). See also Baer, supra note 6, at 53 (criticizing the loss variable for encompassing intended loss).

Overall, Table 1 presents initial descriptive patterns that suggest regressive inequality in financial crime prosecutions. First, readers will notice that fraud makes up more than 80% of financial crime cases across all columns, making up 76.5% of low-level cases (column (2)) and 87.8% of high-level cases (column (5)). The median loss in a financial crime prosecution is just under $50,000, but it is $0 in the lowest quartile and nearly $850,000 in the highest. The median fine in all categories—even the most serious financial crimes—is $0.

Table 1 shows there are differences in the representation of defendants by race, gender, and income levels across the severity distribution. Black defendants and female defendants make up a smaller share of defendants in high-loss cases than in other types of cases. Specifically, Black defendants and female defendants each make up around 30–40% of defendants in low to medium-loss cases, but only around 25% of defendants in high-loss cases. Hispanic defendants are particularly overrepresented in low-loss cases. This could be because around half of Hispanic defendants convicted of financial crimes are not U.S. citizens, and among non-citizen defendants many are convicted of crimes that do not involve a concrete victim, such as making a false statement to federal officials or using a false social security number, as in Case A described in Section I.B.

The pattern is similar for education. Defendants who have not completed high school—who are likely to be those with the fewest resources—appear in low-level cases at much higher rates (28% of defendants) than they appear in high-loss cases (11% of defendants). The pattern for defendants who have college degrees—who are likely to be those with the most resources—is the opposite. College graduates make up 31% of defendants in high-loss cases and just 10% of defendants in low-loss cases.

Overall, Table 1 provides initial descriptive evidence of patterns that this Article explores in the next three subsections. It suggests that people who are likely to have the most advantages—people who are male, White, and have completed college—are more frequently prosecuted for more serious financial crimes than others. The rest of this Part examines inequality in the entire data, over time, and by geography.

Table 1.  Federal Financial Crime Prosecutions, 1994–2019
 

All Financial Crimes

(1)

Low Loss

(2)

Med-Low Loss

(3)

Med-High Loss

(4)

High Loss

(5)

Offense Characteristics
Antitrust0.0020.0030.0030.0030.003
Bribery0.0210.0260.0200.0150.015
Counterfeiting/Forgery0.0830.1890.1030.0490.021
Fraud0.8330.7650.8330.8320.878
Tax Offense0.0610.0180.0440.1050.083
Government Victim0.2460.3050.3150.2940.151
Private Victim0.4220.2980.4400.4550.540
No Concrete Victim0.0570.1420.0350.0200.008
Unknown Victim0.2760.2560.2100.2310.302
Loss (median in $)48,362020,802105,997847,375
Defendant Characteristics
Black0.2930.3020.3840.3240.237
Hispanic0.1470.1990.1140.1150.137
Other Race/Ethnicity0.0680.0640.0560.0580.065
White0.4920.4340.4460.5030.561
Male0.7020.6880.6100.6740.766
Less than HS0.1890.2820.2110.1570.106
HS Only0.3150.3560.3550.3080.248
Some College0.3100.2580.3160.3430.333
College Grad0.1860.1040.1180.1920.313
U.S. Citizen0.6690.6710.7400.7730.797
Retained Counsel0.3370.2180.2640.4080.562
Fines Waived0.8590.8270.9050.9090.920
Case Characteristics
Guidelines Mean (months)28.712.312.622.654.3
Any Incarceration0.5590.4910.4950.7020.863
Sentence (months)16.48.47.214.236.2
Below Guidelines0.4780.2360.4950.6110.599
In-Range0.4990.7300.4840.3690.383
Above Guidelines0.0210.0330.0190.0180.017
Fine (median in $)00000
Restitution (median in $)5,800011,42265,000429,968
Observations276,21043,15143,14643,22643,071
Note: All variables are coded as 0/1 unless otherwise noted. Guidelines and sentence length variables are capped at 470 months—the Commission’s assigned value for life sentences. Many variables are not reported in all years.

C.  Overall Inequality (All Districts, All Years)

This section begins by examining whether one can fairly say the government focuses its financial crime enforcement efforts on “white-collar” crime. It suggests the answer is no. It shows that low-income and Black defendants are disproportionately represented while higher-income and White defendants are underrepresented in federal financial crime cases relative to the U.S. population. It shows that this overrepresentation is particularly stark for Black women, who are underrepresented in federal criminal cases as a whole but overrepresented in financial crime prosecutions.

The Commission data do not provide information about a person’s income or wealth, so Figure 1 uses three proxies for a defendant’s financial means: the level of formal education attained by the defendant, whether the defendant’s fines were waived by the court based on the defendant’s inability to pay them, and whether the defendant retained paid counsel. Appendix Table A.2 presents the same results in table form.

Figure 1.  Proxies for Poverty in Federal Financial Crime Cases
 
Note: Educational attainment is only reported for defendants sentenced in fiscal years 1997 through 2019. Defense counsel type is only reported for defendants sentenced in fiscal years 1994 through 2003. Waived fines are reported for all years (1994 through 2019).

Figure 1 shows the averages for all federal financial crime defendants (dotted columns), for U.S. citizen financial crime defendants (solid columns), and for the U.S. adult population (striped columns). It reports the estimates separately for U.S. citizen-defendants because Census data, which is used to compute the averages across the U.S. adult population, chronically undercounts people who are not U.S. citizens.140U.S. Census Bureau, Counting the Hard to Count in a Census 1, 4 (July 2019) (listing “[m]igrants and minorities” as a population in the U.S. that is “hard-to-count,” which is defined as a population “for whom a real or perceived barrier exists to full and representative inclusion in the [Census] data collection process”). Despite this undercounting, the averages for U.S. citizen-defendants are very similar to the averages among all federal defendants.

As Figure 1 shows, nearly 20% of financial crime defendants did not graduate high school, which is true of only around 10% of U.S. adults. Around 30% of the U.S. adult population has a college degree, but less than 20% of federal financial crime defendants have one. Around 85% of federal financial crime defendants have their fines waived by the court. Put another way, only around 15% of federal financial crime defendants can afford to pay their fines. The majority (around two-thirds) of federal financial crime defendants rely on appointed counsel. These averages suggest that defendants convicted of financial crimes are likely to have a lot less income and wealth than the average U.S. adult.

Figure 2 displays race-gender representation in federal financial crime cases (solid columns) and all federal criminal cases (striped columns) over the years 1994 to 2019. Appendix Table A.3 presents the same results in table form. The horizontal line at y = 1 demarcates the boundary for whether a group is over- or under-represented in federal cases relative to their share of the U.S. adult population.141For each group, the column height represents the share of defendants in that group divided by the share of people in that group in the U.S. adult population over the period 1994–2019. For example, Black men make up roughly 18.8% of fraud defendants and roughly 5.5% of the U.S. adult population, so the height of their solid green column is (18.8/5.5) = 3.42. An alternative way to compute inequality would be to subtract rather than divide each defendant group’s representation from their representation in the U.S. adult population. When computed this way, the inequality patterns are similar but less extreme because the race-gender groups are not equally sized.

Figure 2 shows that, as many readers will already know, Black and Hispanic men are the most overrepresented groups in the federal criminal system (their striped columns extend the highest), while women who are not Black or Hispanic are the most underrepresented groups (their striped columns extend the lowest). Overall, there are five race-gender groups that are underrepresented relative to the adult population: women of all race and ethnicity groups and White men. Men who are not Black or Hispanic are prosecuted at rates closest to parity (their striped columns are the shortest), with White men slightly underrepresented and men who are another race slightly overrepresented.

For financial crimes, the pattern is different in a few notable ways. First, unlike in the entire federal defendant population, Black men and women are the most overrepresented groups in financial crime prosecutions, while White and Hispanic women are the most underrepresented groups. Black men are significantly more overrepresented in financial crime prosecutions than any other group (their solid column is much taller than any other solid column). Men who are not Black are also overrepresented in financial crime cases but to a much lesser extent than Black men.

Black women are overrepresented among financial crime defendants despite being underrepresented in the federal criminal defendant population. Women of all other race and ethnicity groups are underrepresented in financial crime prosecutions, just as they are in all federal prosecutions. These findings suggest that financial crime prosecutions are an important site of racial inequality in the federal criminal system and that this inequality uniquely burdens defendants who are Black.

Figure 2.  Race-Gender Representation in Federal Prosecutions, 1994–2019
 
Note: The y-axis is scaled such that a group that is x times overrepresented will have the same size column as a group that is x times underrepresented. BM=Non-Hispanic Black Men; BW=Non-Hispanic Black Women; HM=Hispanic Men; HW=Hispanic Women; OM=All Other Men (including Alaska Native, American Indian, Asian, Native Hawaiian, and Other Pacific Islander Men); OW=All Other Women (including Alaska Native, American Indian, Asian, Native Hawaiian, and Other Pacific Islander Women); WM=Non-Hispanic White Men; WW=Non-Hispanic White Women.

D.  Inequality in Financial Crime Prosecutions over Time

This section describes how the federal financial criminal caseload has changed over the past quarter century. It shows first that the annual number of financial crime prosecutions remained stable until 2015, when it began to decrease. Second, it shows that the caseload decline in 2015 did not coincide with any noticeable change in the education, gender, or race gaps that persist throughout the period. Third, it shows that since around 2008, Black defendants have been prosecuted at roughly three times the per capita rates that Hispanic, non-Hispanic White, and other defendants have been prosecuted for financial crimes. Beginning in 2008, defendants in all racial or ethnicity groups who are not Black were prosecuted at very similar per capita rates. Fourth, it shows that this race gap is larger but shrinking among female defendants and smaller but more stable among male defendants. Because most of the patterns I documented are stable over time, most of the figures that accompany this section are contained in the Appendix.

Before examining how inequality has changed over time, this section first considers how overall levels of financial crime prosecution have changed since the early 1990s. Appendix Figure A.1 plots the federal government’s criminal caseload for the three most-prosecuted types of crime: drug trafficking (dotted line); immigration (dashed line); and financial crime (solid line). As Figure A.1 reveals, the annual number of federal prosecutions of financial crime remained stable until it began to decline in 2015. On the other hand, financial crime as a share of all federal prosecutions has decreased over a longer period, but this is not due to a significant decrease in the number of financial crime cases; rather, it is a result of a steep rise in immigration-related prosecutions, which dilute financial crime’s share of all federal criminal cases.

It is possible that as the number of financial crime prosecutions decreased beginning in 2015, inequality in who is prosecuted for financial crimes also changed. Appendix Figure A.2 looks for changes in the average education levels of defendants prosecuted for financial crimes, while Figure 3 considers changes in the race and gender composition of the financial criminal defendant population over time. Figure A.2 studies defendants’ educational attainment because education proxies for a defendant’s income, which is not a variable that the Commission data reports.142See supra Section II.C.

Figure A.2 plots financial crime cases prosecuted against defendants who did not graduate high school, graduated high school but did not attend college, attended college but did not earn a bachelor’s degree, and earned a bachelor’s degree. Panel A, which plots the share of defendants in each category, shows that the educational composition of financial crime defendants remained largely stagnant over the 1997 to 2019 period.

Panel A suggests a small increase in the share of financial crime defendants who have attended or completed college and a small decrease among defendants who never attended any college over the same period. However, as Panel B reflects, these changes have not kept pace with the population, which has on average seen increased formal education over time. If anything, the education gap expanded over the period, as Panel B shows. Panel B plots the extent to which defendants in each educational group are over- or under-represented relative to the U.S. adult population. It shows that defendants who have not completed high school were prosecuted at higher rates in the late 2010s than in earlier parts of the period. Thus, Figure A.2 suggests that overall changes in the financial crime caseload over time did not benefit those with few resources; if anything, the opposite is true.

Financial crime also has a race and gender gap. As with nearly all types of crime, men are more likely to be prosecuted for financial crimes than women. Racial gaps in financial crime also persist among both male and female defendants. Figure 3 plots the per capita rates at which each race-gender group is prosecuted for financial crimes over the 1994–2019 period. To avoid cramming eight lines into one graph, Panel A plots the prosecution rates for female defendants and Panel B for male defendants. The panels are arranged side-by-side and scaled with the same y-axis so that readers can compare female and male defendants by looking across the panels. The y-axis measures the number of financial crime defendants in each race-gender group divided by the U.S. adult population of that race-gender group (then multiplied by 1000). Thus, a higher line indicates a higher rate of prosecution.

Figure 3.  Financial Crime Cases by Race, 1994–2019
A.  Female DefendantsB.  Male Defendants
  
Note: Each line represents the number of financial crime cases brought against defendants in the race-gender group, multiplied by 1,000 and divided by the U.S. adult population of that race-gender group. Race-gender groups are labeled as in Figure 2.

Figure 3 shows several facts about race and gender inequality in federal financial crime prosecutions. First, financial crime has a persistent gender gap. Men are prosecuted for financial crimes at higher per capita rates than women. Second, Figure 3 shows that Black men and women are prosecuted for financial crimes at the highest rates. Since 2008, there does not appear to be a significant race gap among any other race groups for either female or male defendants. Instead, Black adults are uniquely susceptible to prosecution for financial crimes.

Third, the racial gaps appear to narrow over time for women but not men. For female defendants, Panel A shows that prosecution rates among racial groups compressed over the 1994 to 2019 period. The data bears this pattern out: Black women comprised 38% of female financial crime defendants in 1994 and 32% in 2019.143This observation is based on the author’s analysis of the data. See Didwania, Data, supra note 17. For male defendants, Panel B shows less compression. The data also bears this pattern out: Black men comprised 25% of male financial crime defendants in 1994 and 27% in 2019.144Id. Over this period, Black men and Black women constituted between 5–7% of the U.S. adult population, so these changes cannot be attributed to significant changes in the composition of the underlying population.145In 1994, Black men made up 5% of the U.S. adult population and Black women made up 6%. In 2019, Black men made up 6% of the U.S. adult population and Black women made up 7%.

E.  Inequality in Financial Crime by Geography

The previous section showed that over the last three decades, financial crime cases have remained a significant portion of the federal criminal docket and that income, gender, and racial inequalities persist in these prosecutions. Among male and female defendants, Non-Hispanic Black people are prosecuted at roughly three times the per capita rate as all other defendants. People who did not complete high school are by far the most overrepresented group in financial crime cases, while those who have completed college are the only group that is significantly underrepresented.

But averages across the entire federal criminal system as presented in the previous sections obscure differences in how individual USAOs prosecute financial crime. For example, the previous sections showed that Black women and Black men are overrepresented in federal financial crime cases while White women are underrepresented, but one might wonder whether this is true in all federal districts in the United States. Variation over the entire country might reflect variation in underlying rates of financial crime, office priorities, or the individual attitudes of decisionmakers such as prosecutors and agents. This section measures and maps inequalities in financial crime prosecutions at the USAO level. Figure 4 begins by showing the intensity with which each USAO prosecutes financial crimes. Darker shading means a larger share of the district’s cases are financial crime cases.

Figure 4 shows that the districts that focus more heavily on fraud cases include large urban districts like the Central District of California (home to Los Angeles), the Northern District of Illinois (home to Chicago), and the Southern District of New York (home to Manhattan). In these USAOs, financial crime respectively constitutes 32.1%, 36.7% and 29.3% percent of all criminal cases. This finding is perhaps unsurprising because these districts encompass many major financial centers. The five districts that border Mexico have much less intense financial crime caseloads (less than 6% of all prosecutions in all five districts) because immigration cases dominate the federal criminal caseloads in those districts.146This observation is based on the author’s analysis of the data. See Didwania, Data, supra note 17. The five federal districts that border Mexico are the District of Arizona, the Southern District of California, the District of New Mexico, the Southern District of Texas, and the Western District of Texas. Together, the USAOs in these five districts prosecuted 32% of all federal criminal cases between 1994 and 2019. Id. In these USAOs, immigration cases made up 55% of the caseload. In the remaining 88 USAOs, immigration cases made up 10% of the caseload. Id. Figure 4 also shows that USAOs in Western states appear to prosecute financial crime less intensely than states in the Deep South147The term “Deep South” does not have a settled definition. Most definitions suggest the core states are Alabama, Georgia, Louisiana, Mississippi, and South Carolina, which is the definition used in this Article. and the Great Lakes Region.148The Great Lakes region includes Illinois, Indiana, Michigan, Minnesota, New York, Ohio, Pennsylvania, and Wisconsin.

Figure 4.  Financial Crime Prosecution Intensity (All Years)
 
Note: This figure maps the share of each district’s criminal cases that are financial crime cases. Each shade represents an equal interval in the distribution. Lightest shading means roughly 2–11% of cases in the district are financial crime cases; second-lightest shading means 11–19% of cases are financial crime cases; second-darkest shading means 19–28% of cases are financial crime cases; and darkest shading means 28–37% of cases are financial crime cases.

Figure A.3 shows that districts in the Deep South, Alaska, and Oklahoma prosecute women for financial crimes at among the highest rates in the United States. Figure A.3 plots the intensity with which each district prosecutes women for financial crimes relative to men. Darker shading means female defendants make up a larger share of that USAO’s financial crime caseload. There are eight districts in which women constitute more than forty percent of financial crime defendants: the Southern, Middle, and Northern Districts of Alabama; the District of Alaska; the Middle District of Georgia; the Middle and Western Districts of Louisiana; and the Northern District of Oklahoma. By contrast, women make up the smallest portion of fraud defendants in New England and southwestern states. There are ten districts in which women make up less than twenty-five percent of fraud defendants: the Southern District of California, the District of Connecticut, the District of Massachusetts, the District of Minnesota, the District of New Hampshire, the District of New Jersey, the Eastern and Southern Districts of New York, the Eastern District of Pennsylvania, and the District of Rhode Island.

Prosecuting women for financial crimes at higher rates in the Deep South, Alaska, and Oklahoma compared with other jurisdictions is likely to create racial inequality among female defendants because the Deep South states have among the largest Black populations in the United States.149Over the 1994–2019 period, the states with the largest Black adult populations were Mississippi (34% of adults); Louisiana (30% of adults); Georgia (28% of adults); Maryland (28% of adults); South Carolina (27% of adults); and Alabama (24% of adults). Alaska and Oklahoma have among the largest Indigenous populations in the United States.

Figure 5 explores the geography of race and gender inequality in financial crime prosecutions. It depicts whether race-gender groups are over- or under-represented in financial crime prosecutions relative to their share of the U.S. adult population in each federal district. In Figure 5, districts filled in blue stripes mean the group is underrepresented (with darker shades of blue representing more underrepresentation). Districts filled in solid red mean the group is overrepresented (with darker shades of red representing more overrepresentation).

Panels A and B show that Black men are overrepresented in financial crime cases in every federal district, and Black women are overrepresented in all but six federal districts.150The six districts in which Black Women are underrepresented in financial crime cases relative to their share of the adult population are the District of Columbia, the Southern District of California, the Southern District of Florida, the District of New Jersey, the Eastern District of New York, and the Southern District of New York. In contrast, Panel H shows that White women are underrepresented in financial crime cases in every federal district. White men are overrepresented in roughly half of all districts, but in all districts, it is clear their representation is relatively close to parity because all of the districts have pale shading. These findings demonstrate that the racial inequalities documented across the full United States are generated at least in part by inequalities within—not just between—USAOs.

As in Figure 5, Figure 6 explores the geography of income inequality in financial crime prosecutions. As throughout, the defendant’s level of formal education is used as a proxy for income because the Commission data does not report information about a defendant’s income or wealth. Also, like Figure 5, Figure 6 uses red solid- blue striped shading to indicate whether defendants are over- or under-represented relative to the U.S. adult population. Districts shaded in blue stripes mean the group is underrepresented (with darker shades of blue representing more underrepresentation). Districts shaded in solid red mean the group is overrepresented (with darker shades of red representing more overrepresentation). The shading in Figure 6 uses the same red/blue scale as Figure 5 so readers can compare.

Figure 5.  Race-Gender Representation in Financial Crime Prosecutions
A.  Cases Against Black MenB.  Cases Against Black Women
  
C.  Cases Against Hispanic MenD.  Cases Against Hispanic Women
  
  
E.  Cases Against Men of Another RaceF.  Cases Against Women of Another Race
  
G.  Cases Against White MenH.  Cases Against White Women
  
Note: This figure maps the over- and under-representation of each race-gender group in the district’s financial crime cases. Districts shaded in striped (solid) fills prosecute the race-gender group at lower (higher) rates than the district’s population. Darker shading indicates larger disparity.
Figure 6.  Educational Representation in Financial Crime Prosecutions
A.  Cases Against Defendants with a College Degree
 
B.  Cases Against Defendants Without a High School Degree
 
Note: This figure maps the over- or under-representation of defendants in the district’s financial crime cases. Districts shaded in striped (solid) fill prosecute the education group at lower (higher) rates than the district’s population. Darker shading indicates larger disparity.

Figure 6 shows that defendants who have graduated from college—and are likely to be the wealthiest federal defendants—are underrepresented in financial crime prosecutions in every federal district in the United States, even those that prosecute the most complex and sophisticated financial crime (such as the Southern District of New York). In contrast, defendants who have not completed high school—and are likely to have the fewest resources—are overrepresented in nearly every district, although they are underrepresented in eleven districts.

The preceding discussion suggests that USAOs could significantly vary in the average severity of financial crimes they prosecute. Figure 7 investigates this theory and depicts the median loss associated with financial crime cases in each federal district. In other words, Figure 7 shows the severity of the average financial crime prosecution by each USAO. It shows significant variation in severity across USAOs.

Figure 7.  Median Loss Amount in Financial Crime Prosecutions (All Years)
 
Note: This figure maps the median loss in financial crime cases by USAO. Each shade represents an equal interval. Lightest shading means the median loss amount in financial crime cases is between $7,519 and $44,323; second-lightest shading means the median loss is between $44,323 and $81,127; second-darkest shading means the median loss is between $81,127 and $117,931; and darkest shading means the median loss is between $117,931 and $154,735.

Figure 7 shows that the most serious financial crimes are prosecuted in the Northeast (including the Eastern and Southern Districts of New York, and the Districts of Connecticut, Massachusetts, and Rhode Island), as well as a few scattered districts that are home to major U.S. cities (the Southern District of California, the Southern District of Florida, the Northern District of Georgia, the Northern District of Illinois, and the District of Minnesota). The least serious financial crimes are prosecuted in Southern and Great Plains states.

III.  EXPLAINING THE FINDINGS

Part II presented evidence of income, racial, and gender inequality in the prosecution of federal financial crimes. It showed that the federal prosecution of financial crime has a disparate impact, prosecuting low-income and Black people at higher rates than the rest of the U.S. adult population, while prosecuting college graduates and White people at lower rates than the rest of the adult population. Part II also showed that these inequality patterns have persisted since the 1990s and appear in every federal judicial district. This Part offers several potential explanations for the inequalities documented in Part II. It first examines differences in the reported offense conduct of financial crime defendants in different education, race, and gender groups. It shows that the defendant groups that are the most overrepresented are also prosecuted for, on average, the least serious financial crimes. It then describes how systemic incentives, formal law and policy, and individual biases could explain the Article’s findings. I do not attempt to definitively prove that any particular mechanism dominates. Instead, this Part is designed to present many possible explanations for the regressive nature of federal white-collar prosecution.

A.  Charged Offense Conduct

As a threshold matter, this section examines whether the federal financial crime cases brought against defendants of different education, race, and gender groups systematically differ in reported offense conduct. The DOJ and FBI routinely state that they prioritize prosecuting serious and sophisticated financial crimes. It could be that the groups that are most overrepresented in financial crime prosecutions also commit on average the most serious financial criminal offenses, and that overrepresentation is thus consistent with the federal government carrying out its stated priorities. This section considers but rejects that hypothesis.

To perform this analysis, this section considers three variables that capture offense conduct: (1) the offense severity (which primarily corresponds with the amount of monetary loss in financial crime cases); (2) whether the case involved illegal drugs; and (3) the average amount of aggravation computed in the case. I define the amount of aggravation as the amount by which the defendant’s base offense level was increased or decreased at sentencing on account of their offense characteristics.151The U.S. Sentencing Guidelines Manual identifies many offense characteristics that can increase or decrease the advisory sentencing range for a person convicted of a financial crime. For example, a person’s offense level will increase if their conduct “resulted in substantial financial hardship” to multiple victims, or if it involved damage to “property from a national cemetery or veterans’ memorial,” or if it involved the misappropriation of a trade secret, among other things, U.S. Sent’g Guidelines Manual §§ 2B1.1(b)(2)(B)–(C), 2B1.1(b)(5), 2B1.1(b)(14) (U.S. Sent’g Comm’n 2021). The aggravation measure can therefore be a positive or negative number. Figure 8 presents averages for each of these measures of offense conduct by defendants’ educational attainment. As before, I use the defendant’s level of formal education as a proxy for income because the Commission data do not include information about defendants’ income or wealth.

Figure 8.  Financial Crime Case Characteristics by Education Group
A.  Median Loss Amount in Financial Crime Cases
 
B.  Share of Financial Crime Cases Involving Drugs
 
C.  Average Aggravation in Financial Crime Cases
 
Note: Average aggregation is the average difference between defendants’ base and final offense levels.

Figure 8 suggests that financial crime defendants who have attained more formal education are prosecuted for financial crimes that are more serious than the financial crimes prosecuted against defendants with less formal education. The median loss amount for defendants without a high school diploma is just $18,500, while the median loss amount for defendants with a college degree is $168,276. The amount of aggravation in the offense is also increasing in formal education, as Panel C shows. In contrast, Panel B shows that the presence of illegal drugs in financial crime cases is roughly equal across all education groups.

Figure 9 plots the same three variables by defendant race-gender group. Figure 9 demonstrates that Black men and women—who Part II showed are prosecuted for financial crimes at the highest rates—do not commit the most serious financial crimes. Cases involving female defendants also tend to be less severe than those against male defendants. Median loss amounts for female financial crime defendants are lower than for male financial crime defendants in all racial groups except Hispanic defendants, in which loss amounts are roughly equal between male and female defendants. In all racial groups, female financial crime defendants are less likely to have drugs involved in their cases. Finally, financial crime cases against women involve fewer aggravating characteristics.

In all measures, financial crime cases brought against White men appear to be the most serious. They involve by far the largest losses—the median loss amount for financial crime prosecutions of White men is $80,150; for Black women and women who are not White, Hispanic, or Black, the amount is $29,520 and $29,416, respectively. Financial crime cases against White men are also the most likely to involve drugs and the largest average aggravation.

Given that differences in offense conduct do not appear to justify the inequalities documented in Part II, the remaining sections explore alternative explanations for the findings. The data do not allow me to disentangle whether the inequalities documented in this Article are created by intentional discrimination, subconscious bias, are a byproduct of systemic incentives that shape prosecutorial and investigative decisions about which cases to prioritize, or are some combination of all these (or other) reasons. Sections III.B, III.C, and III.D consider many explanations for the findings.

Figure 9.  Financial Crime Case Characteristics by Race-Gender Group
A.  Median Loss Amount in Financial Crime Case
 
B.  Share of Financial Crime Cases Involving Drugs
 
C.  Average Aggravation in Financial Crime Cases
 
Note: Race-gender groups are labeled as listed in Figure 2. Average aggregation is the average difference between defendants’ base and final offense levels.

B.  Systemic and Structural Explanations

In many areas of law, the government struggles to aggressively prosecute or pursue legal claims against sophisticated lawbreakers. This section focuses on systemic explanations for why federal prosecutors might focus on lower-level financial crime cases. It argues that complicated financial crimes are difficult to detect, hard to investigate, and burdensome to prove. As Jesse Eisinger put it, “Embezzlement is as easy to understand as purse snatching. But securities manipulation is a more abstract concept.”152Eisinger, supra note 6, at 59. The workplace realities that prosecutors and investigators confront could create the inequalities documented in Part II.

The inequalities in financial crime prosecutions might reflect structural realities that have been documented in many other settings. In an article examining how the federal government prosecutes drug crime, for example, Lauren Ouziel lays bare the “disconnect between [federal criminal] law’s ambition and fruition.”153Ouziel, supra note 78, at 1077. Ouziel shows that in federal drug prosecutions, the substantive criminal law is explicitly designed to target the most serious defendants—those whose crimes involve large quantities of illegal drugs and acts of physical violence, and those who have significant prior criminal records.154See id. at 1079. But despite this ambition, the federal government nonetheless prosecutes many defendants who do not fall into these categories.155See id. Ouziel argues that the pressure and incentives that federal prosecutors face in their work—among other things—contribute to this ambition/fruition divide.156See id. at 1110–11 (arguing that because it is difficult for the federal government to monitor prosecutors’ “performance” in enforcing federal drug laws, it turns to “proxies” such as arrests and seizures).

Examples of the ambition/fruition divide are not limited to the criminal setting. In the context of environmental enforcement, Nathan Atkinson shows that the Environmental Protection Agency (“EPA”) imposes fees on corporate pollution that are roughly one-fifth the size necessary to make polluting unprofitable ex ante.157Nathan Atkinson, Profiting from Pollution, 41 Yale J. Regul. 1, 5–6 (2023); see also Roy Shapira & Luigi Zingales, Is Pollution Value-Maximizing? The Dupont Case 1 (Nat’l Bureau of Econ. Rsch., Working Paper No. 23866, 2017) (showing that DuPont’s toxic pollution—which ultimately led to a roughly one billion-dollar judgment against the company—was a rational, profit-maximizing choice rather than the result of ignorance or poor governance). In another example, ProPublica journalists Paul Kiel and Jesse Eisinger showed a perhaps illogical disparity in the Internal Revenue Service (“IRS”) enforcement efforts: taxpayers who receive the Earned Income Tax Credit (“EITC”)—mostly low-income wage earners—are audited at higher rates than households with much larger earnings.158Paul Kiel & Jesse Eisinger, Who’s More Likely to be Audited: A Person Making $20,000—or $400,000?, ProPublica (Dec. 12, 2018, 5:00 AM), https://www.propublica.org/article/earned-income-tax-credit-irs-audit-working-poor [https://perma.cc/5CF6-YGWB] (showing that in 2017, EITC recipients were audited at twice the rate of taxpayers with incomes between $200,000 and $500,000). Along the same lines, a county-level analysis by ProPublica’s Paul Kiel and Hannah Fresques found that America’s poorest counties are our most audited.159Paul Kiel & Hannah Fresques, Where in the U.S. Are You Most Likely to Be Audited by the IRS?, ProPublica (Apr. 1, 2019), https://projects.propublica.org/graphics/eitc-audit [https://perma.cc/DH7Q-ER5A]. Yet recent research shows that despite the lower costs to carry them out, IRS audits of low-income people yield less net revenue than audits of wealthy taxpayers at the top of the income distribution.160William C. Boning, Nathaniel Hendren, Ben Sprung-Keyser & Ellen Stuart, A Welfare Analysis of Tax Audits Across the Income Distribution 1 (Nat’l Bureau of Econ. Rsch., Working Paper No. 31376, 2023). IRS’s choice to focus much of its enforcement activity on EITC filers also contributes to racial inequality in audits.161This is because Black taxpayers are more likely to claim the EITC than non-Black taxpayers, EITC claimants are audited at high rates, and because among EITC recipients, Black taxpayers are more likely to be audited than non-Black taxpayers. See Hadi Elzayn, Evelyn Smith, Thomas Hertz, Arun Ramesh, Robin Fisher, Daniel E. Ho & Jacob Goldin, Measuring and Mitigating Racial Disparities in Tax Audits 3–4 (Stanford Inst. for Econ. Pol’y Rsch., Working Paper, 2023), https://dho.stanford.edu/
wp-content/uploads/IRS_Disparities.pdf [https://perma.cc/D8QN-W35Y] (analyzing around 150 million tax returns and estimating that Black taxpayers are audited at higher rates than non-Black taxpayers and that this difference is primarily driven by the difference in audit rates among taxpayers who claim the EITC). See generally Jeremy Bearer-Friend, Colorblind Tax Enforcement, 97 N.Y.U. L. Rev. 1 (2022) (arguing that IRS enforcement decisions are vulnerable to racial bias even though the IRS does not ask taxpayers to identify their race or ethnicity when they file tax returns).

Like the drug crime and IRS contexts, prosecutors and law enforcement agents working on financial crimes face incentives and constraints that likely lead them to focus their efforts on straightforward, uncomplicated, and winnable prosecutions.162See Stuntz, supra note 72, at 535 (“[Unelected line prosecutors] are likely to seek to make their jobs easier, to reduce or limit their workload where possible. That inclination means two things: limiting the number of cases on their dockets, and limiting the cost of the process per case.” (citation omitted)). Of course, what kinds of cases and defendants an agent or prosecutor thinks are “winnable” requires judgments that will be filtered through and reinforced by the agent or prosecutor’s individual biases, as described in Section III.D.

How do prosecutors decide which potential cases are winnable? They likely consider the evidentiary strength of their case, the resources necessary to investigate and prosecute the case, and how a jury is likely to view the case.163See Anna Offit, Prosecuting in the Shadow of the Jury, 113 Nw. U. L. Rev. 1071 (2019) (presenting ethnographic research showing that federal prosecutors think about how hypothetical jurors will view their cases when making investigative and plea bargaining decisions). These assessments are likely shaped by biases, as described in the next section.

All these factors—the strength of the evidence, the resources necessary to bring the case, and how a jury is likely to view the case—militate toward prosecuting low-level cases. As described in Sections I.B and III.C, a financial crime prosecution typically requires a prosecutor to prove beyond a reasonable doubt that the defendant intended to defraud someone. In simplistic cases, such as when an employee uses a company credit card to buy personal items, the evidence of fraud will often be straightforward and easily attainable: typically, the victim (the employer) will have records showing unauthorized purchases and can turn those records over to prosecutors.

In contrast, the task of building a case will be much more difficult in frauds for which there is no victim who can provide evidence of the fraud, such as when a fraud is carried out in a large corporate organization with many diffuse victims. As Miriam Baer describes, “[l]ife within corporate settings is remarkably compartmentalized and siloed. Information and responsibility fractures among multiple units and departments, allowing criminal targets to claim that the left hand did not know what the right hand was doing, or at very least, that an intent to harm or deceive was absent.”164Baer, supra note 6, at 110.

In such cases, the government will typically need to rely on a whistleblower for evidence and may have a hard time proving that any particular person involved had the requisite intent to defraud. Whistleblowers can be hard to recruit because, although they are occasionally rewarded for bringing wrongdoing to light, more often they are fired and struggle to find a new job in their industry.165William D. Cohan, High Risk but Little Reward for Whistle-Blowers, N.Y. Times (Mar. 26, 2015), https://www.nytimes.com/2015/03/27/business/dealbook/high-risk-but-little-reward-for-whistle-blowers.html [https://perma.cc/22RX-N2PG]; see also William D. Cohan, Wall St. Whistle-Blowers, Often Scorned, Get New Support, N.Y. Times (Feb. 11, 2016), https://www.nytimes.com/2016/02/12/business/dealbook/wall-st-whistle-blowers-often-scorned-get-new-support.html [https://perma.cc/ZHR4-XUYS] (describing an advocacy group, Bank Whistleblowers United, “that aims to improve the status of Wall Street whistle-blowers and change the way Wall Street is regulated”); Alexander I. Platt, The Whistleblower Industrial Complex, 40 Yale J. Regul. 688, 707–09 (2023). This is precisely what happened to Alayne Fleischmann, the whistleblower in Case D.166See Daniel C. Richman, Corporate Headhunting, 8 Harv. L. & Pol’y Rev. 265, 269 (2014) (describing likely difficulties in bringing criminal charges against individuals involved in the 2008 financial crisis). But see Miriam Baer, supra note, 6, at 15 (“[W]hite-collar crimes are not always as difficult to prove as some commentators suggest . . . . When the government feels like it, it mobilizes its extensive resources.”).

Second, building and bringing complex cases takes a lot of work and resources. It uses up prosecutors’ and investigators’ time. The more witnesses there are to interview, the more documents there are to review, and the more expertise is required to understand the fraud—all these tasks require a lot of resources. A straightforward case can move forward more quickly and easily.

Relatedly, the resource differences on each side of a criminal case can strain the government’s ability to prosecute. Charging a person who will hire a large law firm to represent them in defense will create a different resource dynamic than prosecuting a person who will rely on appointed counsel.167Of course, there are many talented attorneys who work as appointed counsel, but they do not have the same level of resources as a large law firm. Some research has found that attorneys who are retained rather than appointed appear to achieve better outcomes for their clients. See, e.g., Amanda Agan, Matthew Freedman & Emily Owens, Is Your Lawyer a Lemon? Incentives and Selection in the Public Provision of Criminal Defense, 103 Rev. Econ. & Stat. 294, 294 (2021) (finding worse outcomes for criminal defendants represented by appointed rather than retained counsel); Thomas H. Cohen, Who is Better at Defending Criminals? Does Type of Defense Attorney Matter in Terms of Producing Favorable Case Outcomes, 25 Crim. J. Pol’y Rev. 29, 29 (2014). Several studies also show that federal public defenders outperform Criminal Justice Act panel attorneys. Radha Iyengar, An Analysis of the Performance of Federal Indigent Defense Counsel 2 (Nat’l Bureau of Econ. Rsch., Working Paper No. 13187, 2007); see also Michael A. Roach, Indigent Defense Counsel, Attorney Quality, and Defendant Outcomes, 16 Am. L. & Econ. Rev. 577, 615 (2014). These resource differences could easily lead the federal government to disproportionately prosecute indigent defendants.

C.  Formal Law and Policy

The substantive laws and rules that define financial crimes and govern how they are prosecuted and sentenced favor sophisticated criminal lawbreakers in many ways. We see examples of this phenomenon in other contexts, too. For example, by far the largest source of theft in the United States is wage theft, which some researchers estimate accounts for more than $15 billion stolen every year.168David Cooper & Teresa Kroeger, Employers Steal Billions from Workers’ Paychecks Each Year, Econ. Pol’y Inst. (May 10, 2017), https://www.epi.org/publication/employers-steal-billions-from-workers-paychecks-each-year [https://perma.cc/K74Q-7Q92]. An employer commits wage theft when they do not pay an employee wages to which the employee is legally entitled, such as by paying less than the minimum wage, not paying required overtime wages, or asking employees to work “off the clock” before or after their shifts.169Ihna Mangundayao, Celine McNicholas, Margaret Poydock & Ali Sait, More than $3 Billion in Stolen Wages Recovered for Workers Between 2017 and 2020, Econ. Pol’y Inst. (Dec. 22, 2021), https://www.epi.org/publication/wage-theft-2021 [https://perma.cc/R7W4-ZBVY]. For a comprehensive examination of efforts to criminalize wage theft, see generally Levin, supra note 111. But wage theft is almost never prosecuted.170See Chris Opfer, Prosecutors Treating ‘Wage Theft’ as a Crime in These States, Bloomberg L. (June 26, 2018, 3:31 AM), https://news.bloomberglaw.com/daily-labor-report/prosecutors-treating-wage-theft-as-a-crime-in-these-states [https://perma.cc/4QSZ-RKX8] (noting that “[w]hen a business doesn’t pay workers minimum wages or overtime, it usually risks a government investigation or private lawsuit,” but that “[p]rosecutors in New York and California are starting to view wage violations as an actual crime more often, as opposed to a matter for civil courts”). The primary way that stolen wages are recovered is through civil actions brought by the U.S. Department of Labor’s Wage and Hour Division, state departments of labor, state attorneys general, and civil class actions. In contrast, larceny and auto theft each steal around $5 billion per year and robbery steals around $380 million.171Table 23: Offense Analysis, Number and Percent Change, 2018–2019, U.S. Dep’t of Just., Fed. Bureau of Investigation, 2019 Crime in the United States, https://ucr.fbi.gov/crime-in-the-u.s/2019/crime-in-the-u.s.-2019/tables/table-23 [https://perma.cc/2UTR-ZPPV]. Unlike wage theft, these crimes are frequently prosecuted.172According to FBI statistics, police clear around thirty-one percent of robberies, fourteen percent of auto thefts, and eighteen percent of larceny offenses. Table 25: Percent of Offenses Cleared by Arrest or Exceptional Means, by Population Group, 2019, U.S. Dep’t of Just., Fed. Bureau of Investigation, 2019 Crime in the United States, https://ucr.fbi.gov/crime-in-the-u.s/2019/crime-in-the-u.s.-2019/topic-pages/tables/table-25 [https://perma.cc/SNG9-GJNQ].

There are myriad ways that federal criminal law and formal policy similarly benefit sophisticated people who commit higher-value, more complex crimes. Here, I focus on two: the mens rea requirements of fraud statutes, and the way restitution is calculated and prioritized.

1.  Mens Rea Elements

As described in Section I.B, most financial crimes contain mens rea elements that require the government to prove the defendant’s intent to defraud. In a relatively straightforward fraud—such as Cases A, B, and C described in Section I.C—it is easy to see how a jury could view the defendants’ conduct and conclude that they intentionally deceived their victims. But in a complex fraud case involving many parties, such as Case D, proving a deceitful intent or scheme on the part of any particular participant could be very difficult for prosecutors.173Daniel Richman is more skeptical of claims that proving criminal intent is a significant hurdle to white-collar prosecutions in the context of the financial crisis, noting that mens rea elements “are far from trivial burdens, but prosecutors regularly meet them in any number of mundane white-collar cases.” Richman, supra note 166; see, e.g., Danielle Kurtzleben, Too Big to Jail: Why the Government Is Quick to Fine but Slow to Prosecute Big Corporations, Vox (July 13, 2015, 10:52 AM), https://www.vox.com/2014/11/16/7223367/corporate-prosecution-wall-street [https://perma.cc/N4AM-H57C] (quoting Brandon Garrett as explaining that in the aftermath of the 2008 financial crisis, prosecutors preferred to focus on “crimes that seem tangential to the crisis . . . . where it [was] easier to show that a small number of people had intent . . . versus some of the mortgage fraud, where there [were] sophisticated actors working with each other, where to show intent to defraud [prosecutors would] have to show that there [was] a clearly deceptive scheme that misled someone else”). As a result, complicated and sophisticated financial crimes—which Table 1 shows are more likely to be perpetrated by people who are high-income, male, and White—are likely much more difficult to prosecute.

2.  Restitution Calculations

The rules around restitution calculations also benefit defendants who commit complex crimes. As described in Section I.B, federal law (like the law in all states) requires courts to order restitution in any case “in which an identifiable victim or victims has suffered a physical injury or pecuniary loss.”17418 U.S.C. §§ 3663A(a)(1), 3663A(c)(1)(B).

One might imagine this means people who commit more complex, higher-value crimes will have to pay more restitution and could therefore be more desirable to prosecute from a prosecutor’s perspective. But this is not the case because the restitution statute contains two exceptions. First, it does not require restitution in cases in which “the number of identifiable victims is so large as to make restitution impracticable.”175Id. § 3663A(c)(3)(A). Second, it does not require restitution in cases in which “determining complex issues of fact related to the cause or amount of the victim’s losses would complicate or prolong the sentencing process to a degree that the need to provide restitution to any victim is outweighed by the burden on the sentencing process.”176Id. § 3663A(c)(3)(B). In other words, financial crimes that are more complex, for which losses are harder to calculate, and for which there are more victims are much less likely to involve restitution. Thus, even if JPMorgan Chase or any of its employees had been convicted of a crime in connection with the financial crisis, they would have had a strong argument that the statute did not require them to pay restitution. In contrast, the defendants in Cases B and C were ordered to pay restitution because their crimes were not complex enough to trigger a statutory exception.

3.  Restitution Policy

Notwithstanding the statutory exceptions, federal prosecutors and judges tend to be highly committed to ensuring as much restitution as possible for victims of financial crimes. For example, the federal sentencing statute instructs judges to consider “the need to provide restitution to any victims of the offense” when sentencing defendants.177Id. § 3553(a)(7). The Justice Manual tells prosecutors that when “determining whether it would be appropriate to enter into a plea agreement,” they should consider (among other factors) “[t]he interests of the victim, including any effect upon the victim’s right to restitution.”178U.S. Dep’t of Just., supra note 73, at § 9-27.420. Similarly, the Manual instructs prosecutors to “take[] into account the need for the defendant to provide restitution to any victims of the offense” when making sentencing recommendations.179Id. at § 9-27.730. Assistant Attorney General for the Criminal Division Kenneth A. Polite, Jr. described federal white-collar efforts in a recent speech, telling the audience, “[c]onsidering victims must be at the center of our white-collar cases. . . . Though we cannot always recover every cent, we deploy all tools at our disposal to restrain assets, obtain restitution, and when possible, repatriate assets for victims.”180Polite, supra note 23.

One consequence of prosecutors’ and judges’ desire to provide restitution to victims of financial crimes is that defendants with more resources can argue (either as a pitch to prosecutors before charging or to a judge at sentencing) that they should not be prosecuted or incarcerated because a criminal case or prison sentence will interrupt their ability to earn income to pay toward restitution. For example, a financial advisor convicted of fraud in the District of Massachusetts made this argument in his sentencing memo, writing:

If incarcerated, [the defendant] will not be able to contribute to restitution; he will lose his job and have to start all over upon his release. Whereas in his current position, where he has advanced to a management position in a relatively short amount of time, he will be able to contribute immediately toward a restitution award.181Def.’s Sentencing Mem. at 9, United States v. Cody, No. 17-CR-10291 (D. Mass. Mar. 9, 2019); see also, e.g., Def.’s Sentencing Mem. at 2, United States v. Luna, No. 19 CR 902-1 (N.D. Ill. Nov. 11, 2020) (noting that the defendant already paid some restitution to the victim, was working full-time in a new job and wanted to continue to repay the victim, and arguing that “paying the victim back is a goal the Court should consider in fashioning a non-custodial sentence” for the defendant).

Indeed, federal courts routinely justify low or probation-only sentences for financial crime defendants by stating their desire to allow the defendant to work and provide restitution.182See United States v. Menyweather, 447 F.3d 625, 634 (9th Cir. 2006) (affirming a probation-only sentence for a defendant convicted of fraud and observing “that the district court’s goal of obtaining restitution for the victims of Defendant’s offense . . . is better served by a non-incarcerated and employed defendant”); United States v. Bortnick, No. 03-CR-0414, 2006 U.S. Dist. LEXIS 11744, at *14, *19 (E.D. Pa. Mar. 15, 2006) (imposing a seven-day sentence to a defendant in an $8 million fraud case with a 51–63 month advisory Guidelines range because “[d]efendant owes a substantial amount of restitution, which he will be able to pay more easily if he is not subjected to a lengthy incarceration period”); United States v. Peterson, 363 F.Supp.2d 1060, 1063 (E.D. Wis. 2005) (imposing a one-day sentence so defendant would not lose his job and could pay restitution to the bank he defrauded). But see United States v. Mueffelman, 470 F.3d 33, 40 (1st Cir. 2006) (affirming a 27-month sentence despite the defendant’s argument that “anything beyond a probationary sentence would impair his ability to provide restitution for victims” and his promise to “earn $120,000–175,000 per year to pay toward restitution, with a friend promising to make up any short fall”). In one of the Yale Studies that surveyed federal district court judges about how they sentence white-collar defendants, one judge was asked about his decision not to impose a prison sentence on a person convicted of not reporting large amounts of income. The interviewer asked the judge, “[Y]ou must have considered sending him to a term in prison. What made you decide that that wasn’t appropriate in this case?” The judge responded,

Well, the restitution. There is half a million dollars back in the coffers that we wouldn’t have got if I had sent him to prison. He would have served his term, and there would have been no way of getting it, and eventually some day or other he would have gotten out of the country somehow or other and gotten that money. That was it.183Mann et al., supra note 51, at 492.

A defendant with fewer resources or without stable employment will have a harder time making this argument to a prosecutor, which could explain why wealthy defendants are less likely to be prosecuted for financial crimes.184Indeed, federal prosecutors often decline or defer prosecution of corporations for this reason. See supra notes 85, 109, 110 and accompanying text.

D. Bias

As described in Section I.B, federal investigative agencies and DOJ have nearly absolute discretion in deciding which cases to investigate and prosecute. Although individual agents and federal prosecutors might be constrained formally and informally by office policies and norms, there are almost no formal legal constraints on how enforcement agents decide which cases to investigate and how prosecutors decide which cases to pursue.185See supra note 71 and accompanying text. Wide discretion often allows decisionmakers to make discriminatory decisions, either consciously or subconsciously.

1. Stereotypes About Dishonesty

Deceit is the central characteristic of financial crime. Social psychologists have documented consistent stereotypes that associate honesty with social class, race, and gender in the United States. For example, literature in psychology finds that participants often view people of low socioeconomic status as lazy, incompetent, and prone to substance abuse, while viewing people of high socioeconomic status as more competent and intelligent.186Federica Durante & Susan T. Fiske, How Social-Class Stereotypes Maintain Inequality, 18 Current Op. Psych. 43, 43 (2017).

Stereotypes characterizing women—and, in particular, women of color—as dishonest are pervasive in the United States, which might explain why Black women are overrepresented among financial crime defendants despite being underrepresented in federal prosecutions overall. Women have long been viewed as dishonest in criminal cases,187See, e.g., Diana L. Payne, Kimberly A. Lonsway & Louise F. Fitzgerald, Rape Myth Acceptance: Exploration of Its Structure and Its Measurement Using the Illinois Rape Myth Acceptance Scale, 33 J. Rsch. Personality 27 (1999). and Marilyn Yarbrough and Crystal Bennett describe “a hierarchy when credibility issues arise in the courts. It is not only a simple hierarchy of men over women, but it is one where White women are found to be more credible than African American women.”188Marilyn Yarbrough & Crystal Bennett, Cassandra and the “Sistahs”: The Peculiar Treatment of African American Women in the Myth of Women as Liars, 3 J. Gender Race & Just. 625, 634 (2000) (citing Rosemary C. Hunter, Gender in Evidence: Masculine Norms vs. Feminist Reforms, 19 Harv. Women’s L.J. 127, 165 (1996)). The rhetoric and law of welfare reform in the 1990s also surfaced and magnified already prevalent gender- and race-based stereotypes about dishonesty. Gustafson, supra note 14, at 1 (“[W]hile welfare use has always carried the stigma of poverty, it now also bears the stigma of criminality.”); see also Julilly Kohler-Hausmann, Welfare Crises, Penal Solutions, and the Origins of the “Welfare Queen,” 41 J. Urb. Hist. 756, 757 (2015) (arguing that “opponents of welfare programs recruited the penal system to discredit public aid beneficiaries and administration”); Franklin D. Gilliam, Jr., The “Welfare Queen” Experiment: How Viewers React to Images of African-American Mothers on Welfare, Nieman Reports (June 15, 1999), https://niemanreports.org/articles/the-welfare-queen-experiment [https://perma.cc/3EX2-FLW3] (finding that when White subjects viewed a television story about welfare reform, they were more likely to believe that “welfare recipients cheat and defraud the system” when exposed to a segment that depicted a female benefits recipient as Black compared to one that depicted the female benefits recipient as White). And as Chan Tov McNamarah explains, “[S]kepticism of Black credibility is part of a larger, historically created space in which those who are deemed rational, reliable, and worthy of belief are White and male.”189Chan Tov McNamarah, White Caller Crime: Racialized Police Communication and Existing While Black, 24 Mich. J. Race & L. 335, 372 (2019) (citing Sheri Lynn Johnson, The Color of Truth: Race and the Assessment of Credibility, 1 Mich. J. Race & L. 261 (1996)); see also Kurtis Haut, Caleb Wohn, Victor Antony, Aidan Goldfarb, Melissa Welsh, Dillanie Sumanthiran, Ji-ze Jang, Md. Rafayet Ali & Ehsan Hoque, Could You Become More Credible by Being White? Assessing Impact of Race on Credibility with Deepfakes, ArXiv, Feb. 16, 2021, at 1, 1–2, https://arxiv.org/pdf/2102.08054.pdf [https://perma.cc/E9BJ-XQUG] (displaying Deepfake still photos and video clips that used the same audio but altered the speaker’s race and finding that speaker race had a negligible effect on credibility when presented as a static image but a statistically significant effect when presented as a video (with a White speaker viewed as more credible than a South Asian speaker)). These kinds of prejudices could affect how agents decide which people to investigate and prosecutors decide which cases to bring.

2.  In-Group Favoritism

Bennett Capers argues, “[T]o understand mass incarceration, we must not only understand overcriminalization and overenforcement in minority communities. We must also understand the role played by under-enforcement, and privilege, in nonminority communities.”190I. Bennett Capers, The Under-Policed, 51 Wake Forest L. Rev. 589, 609 (2016). Consciously or not, prosecutors and agents might be less willing to prosecute people with whom they have more in common, a phenomenon often referred to as “in-group favoritism.”

In-group favoritism occurs when a decision-maker gives preferential treatment to those who share a salient trait with the decision-maker, such as by being a member of their gender, racial, ethnic, or religious group.191Jim A.C. Everett, Nadira S. Faber & Molly Crockett, Preferences and Beliefs in Ingroup Favoritism, Frontiers Behav. Neuroscience, Feb. 13, 2015, at 1. In this subsection, I do not mean to rule out that conscious class-, gender-, or race-based bias is also a potential cause of the inequalities documented in Part II. For many years, there was a growing consensus that the majority of discrimination in the United States takes the form of in-group favoritism,192See, e.g., Anthony G. Greenwald & Thomas F. Pettigrew, With Malice Toward None and Charity for Some: Ingroup Favoritism Enables Discrimination, 69 Am. Psych. 669, 669 (2014); Linda Hamilton Krieger, Civil Rights Perestroika: Intergroup Relations After Affirmative Action, 86 Calif. L. Rev. 125 (1998). although in recent years overt racism and sexism have grown increasingly prevalent.193See, e.g., Charles R. Lawrence III, Implicit Bias in the Age of Trump, 133 Harv. L. Rev. 2304, 2311 (2020) (reviewing Jennifer L. Eberhardt, Biased: Uncovering the Hidden Prejudice that Shapes What We See, Think, and Do (2019)) (reflecting on the choice to review “a book about hidden bias when the active threat is self-proclaimed racists marching in the streets[] . . . . [and] when the President of the country was holding rallies and building walls to proclaim himself the protector of a white nation”); see also Griffin Edwards & Stephen Rushin, The Effect of President Trump’s Election on Hate Crimes (Jan. 2019) (working paper), https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3102652 [https://perma.cc/2J38-N6P6].

In-group favoritism is well-documented in the criminal system. In prior work, I showed that federal prosecutors exhibit gender-based in-group favoritism, treating defendants of their own gender relatively more leniently than other-gender defendants.194Stephanie Holmes Didwania, Gender Favoritism Among Criminal Prosecutors, 65 J.L. & Econ. 77, 77 (2022). CarlyWill Sloan has also shown that state-level prosecutors demonstrate race-based favoritism in prosecuting property crimes in New York County. CarlyWill Sloan, Racial Bias by Prosecutors: Evidence from Random Assignment (Jan. 10, 2022) (working paper), https://github.com/carlywillsloan/Prosecutors/blob/master/sloan_pros.pdf [https://perma.cc/7AZT-SF99]. New research suggests that firms risking prosecution appear to strategically leverage in-group favoritism to help improve negotiations with federal prosecutors.195Brian D. Feinstein, William R. Heaston & Guilherme Siqueira de Carvalho, In-Group Favoritism as Legal Strategy: Evidence from FCPA Settlements, 60 Am. Bus. L.J. 5 (2023). Other scholars have previously documented in-group favoritism among other actors in criminal legal systems, including judges196See, e.g., David S. Abrams, Marianne Bertrand & Sendhil Mullainathan, 41 J. Legal Stud. 347, 350 (2012) (finding that African American judges exhibit smaller racial disparities in sentencing than their White counterparts); Oren Gazal-Ayal & Raanan Sulitzeanu-Kenan, Let My People Go: Ethnic In-Group Bias in Judicial Decisions—Evidence from a Randomized Natural Experiment, 7 J. Empirical Legal Stud. 403, 403, 421 (2010) (finding that Arab and Jewish judges in Israel are less likely to detain defendants who share their ethnicity). But see Briggs Depew, Ozkan Eren & Naci Mocan, Judges, Juveniles, and In-Group Bias, 60 J.L. & Econ. 209, 209 (2017) (finding that judges exhibit “negative in-group bias” toward juvenile defendants of the judge’s race); Claire S.H. Lim, Bernardo S. Silveira & James M. Snyder, Jr., Do Judges’ Characteristics Matter? Ethnicity, Gender, and Partisanship in Texas State Trial Courts, 18 Am. L. & Econ. Rev. 302, 305 (2016) (finding that “matches between judges’ and defendants’ ethnicity, race, and gender . . . have negligible effects” on sentence length). and police officers.197See, e.g., Bocar A. Ba, Dean Knox, Jonathan Mummolo & Roman Rivera, The Role of Officer Race and Gender in Police-Civilian Interactions in Chicago, 371 Science 696, 696 (2021) (showing that “Hispanic and Black officers make far fewer stops and arrests and they use force less [often than White officers], especially against Black civilians”); John J. Donohue, III & Steven D. Levitt, The Impact of Race on Policing and Arrests, 44 J.L. & Econ. 367, 367 (2001) (finding that police departments with more minority officers are more likely to arrest White suspects, with little impact on the arrests of non-White suspects); Mark Hoekstra & CarlyWill Sloan, Does Race Matter for Police Use of Force? Evidence from 911 Calls, 112 Am. Econ. Rev. 827, 827 (2022) (finding that “White officers increase force much more than minority officers when dispatched to more minority neighborhoods”). As an important caveat, however, some research finds evidence of a phenomenon called the black-sheep effect, in which people punish in-group members more harshly than out-group members for bad behavior.198See José M. Marques, Vincent Y. Yzerbyt & Jacques-Philippe Leyens, The “Black Sheep Effect”: Extremity of Judgments Towards Ingroup Members as a Function of Group Identification, 18 Eur. J. Soc. Psych. 1 (1988); see also Depew et al., supra note 196, at 233 (finding in-group disfavoritism on the basis of race in juvenile sentencing).

Perhaps more than in other types of federal cases (most of which involve immigration, drugs, or firearm possession), prosecutors and federal agents might feel affinity for financial crime defendants who work as business professionals due to cultural or social proximity. This hypothesis is not new. Over 40 years ago, one of the Yale Studies described in Section I.B.2 surveyed federal district judges and found sentiment of in-group favoritism when judges were asked about sentencing white-collar defendants. For example, one federal judge described his views on sentencing white-collar defendants to prison this way:

I think the first sentence to a prison term for a person who up to now has lived and has surrounded himself with a family, that lives in terms of great respectability and community respect and so on, whether one likes to say this or not I think a term of imprisonment for such a person is probably a harsher, more painful sanction than it is for someone who grows up somewhere where people are always in and out of prison. There may be something racist about saying that, but I am saying what I think is true or perhaps needs to be laid out on the table and faced.199Mann et al., supra note 51, at 486–87.

The authors believe the judge’s previous comment is the result of increased empathy toward wealthy and professional class white-collar defendants.200Id. at 500.

The [judges’] interview responses repeatedly give evidence of the judges’ understanding, indeed sympathy, for the person whose position in society may be very much like their own. In places, the interviews exude the pain that judges feel in seeing the offender uprooted from his family, humiliated before his friends, and exposed to the degradation of imprisonment.

Id.; see also Bibas, supra note 80 (“[J]udges may prefer to look ex post at the sympathetic, white, educated offender who reminds judges of themselves and seems to pose no danger.”).
Indeed, in-group favoritism often takes the form of empathy toward in-group members, and, in experimental settings, people are often more likely to feel empathy in observing the pain of an in-group member compared to an out-group member.201See Mina Cikara, Emile G. Bruneau & Rebecca R. Saxe, Us and Them: Intergroup Failures of Empathy, 20 Current Directions in Psych. Sci. 149, 149 (2011); Jennifer N. Gutsell & Michael Inzlicht, Intergroup Differences in the Sharing of Emotive States: Neural Evidence of an Empathy Gap, 7 Soc. Cognition & Affective Neuroscience 596, 596 (2012); Xiaojing Xu, Xiangyu Zuo, Xiaoying Wang & Shihui Han, Do You Feel My Pain? Racial Group Membership Modulates Empathic Neural Responses, 29 J. Neuroscience 8525, 8525 (2009). It is plausible that prosecutors and FBI agents are more empathetic about the harms of federal prosecution when it comes to potential defendants with similar levels of formal education and wealth.

CONCLUSION

This Article has shown that, contrary to popular wisdom, financial crime is frequently prosecuted in the United States. Part II showed that federal financial crimes are prosecuted in ways that replicate inequalities that exist throughout American criminal law. Black men and women are more likely to be prosecuted for financial crimes than any other racial and gender group. Unlike the traditional view of white-collar crime, which posits that it is a form of crime largely perpetuated by economic elites, the findings also show that federal financial crime defendants are likely to have fewer resources than most U.S. adults.

Part III offered many explanations for these findings. It argued that systemic incentives, formal law and policy, and individual biases could all drive inequality. It also showed that the overrepresentation of Black and low-income defendants does not appear to be because these defendants commit the most egregious forms of financial crime (in fact, the opposite is true).

The inequalities documented in this paper are concerning because they seem to be overlooked. The intense focus on elite white-collar criminals—by the media, the academy, and the federal government itself—seems to at best not understand the realities of the system in which they are operating. This Article hopes to address this mistake.

APPENDIX

Figure A.1.  Federal Criminal Cases: Three Most Common Offense Types, 1994–2019
 
Note: This figure plots the number of cases sentenced each fiscal year between 1994 and 2019 for the three most commonly prosecuted types of federal crime: drug trafficking and possession, immigration, and financial crime.
Figure A.2.  Educational Attainment in Federal Financial Crime Cases Over Time
A.  Educational Attainment in Federal Financial Crime Cases
 
B.  Educational Attainment Representation in Federal Financial Crime Cases
 
Note: For each year, the “Representation Gap” in panel B is computed as the share of financial crime defendants in the educational group divided by the share of the U.S. adult population between the ages of 25 and 54 in that educational group.
Figure A.3.  Gender Inequality in Financial Crime Prosecutions (All Years) 
  
Note: This figure maps the share of each district’s financial crime cases that are prosecuted against women. Each shade represents an equal interval in the distribution. Lightest shading means roughly 15–22% of financial crime defendants in the district are women; second-lightest shading means 22–29% of financial crime defendants are women; second-darkest shading means 29–36% of financial crime defendants are women; and darkest shading means 36–43% of financial crime defendants are women. 
Table A.1.  Victim Coding: Most Prosecuted Financial Crimes
Crime (Short Description)StatuteShare of CasesVictim
Conspiracy or Defrauding the United States18 U.S.C. § 3710.185U
Embezzlement or Theft of Public Money18 U.S.C. § 6410.079G
Attempt or Conspiracy to §§ 1341‑4818 U.S.C. § 13490.074P
Bank fraud18 U.S.C. § 13440.073P
Wire fraud18 U.S.C. § 13430.072P
Mail fraud18 U.S.C. § 13410.065P
Tax Fraud26 U.S.C. § 72010.053G
False Statements to Federal Officials18 U.S.C. § 10010.050N
Counterfeiting18 U.S.C. § 4720.046P
Credit Card Fraud18 U.S.C. § 10290.046P
Identity Theft18 U.S.C. § 10280.044U
Mail Theft18 U.S.C. § 17080.031G
Accessory to a Crime18 U.S.C. § 20.027U
Social Security Fraud42 U.S.C. § 4080.021G
Embezzlement by Bank Employee18 U.S.C. § 6560.015P
Healthcare Fraud18 U.S.C. § 13470.015U
Conspiracy to Defraud the Government18 U.S.C. § 2860.013G
Note: This table reports how victim status was coded for the most prosecuted federal financial crimes. G=government victim; N=no victim; P=private victim; U=unknown victim. The table is restricted to crimes constituting at least one percent of charged cases. Many additional types of financial crimes were also coded, and a complete crosswalk is available from the author by request.
Table A.2.  Proxies for Poverty in Federal Fraud Prosecutions
 

% of Financial Crime Defs

(All)

% of Financial Crime Defs

(Citizens)

% of U.S. Adult Pop

(if applicable)

Less than HS18.8916.7811.09
High School Only31.4931.5929.43
Some College31.0232.2427.42
College Graduate18.6019.4032.06
Fines Waived85.8989.04
Retained Counsel33.7333.63
Observations276,210161,552 
Note: Computations are for federal defendants sentenced under the U.S. Sentencing Guidelines for financial crimes in fiscal years 1994–2019. U.S. adult population averages computed over the years 1994–2019.
Table A.3.  Race-Gender Representation in Federal Fraud Prosecutions
 % of Financial Crime Defs% of All Defs% of U.S. Adult Pop
Black Men18.6219.515.55
Hispanic Men10.8640.377.11
Another Race Men4.693.762.93
White Men36.0622.8232.89
All Men70.2386.4748.47
Black Women10.703.346.41
Hispanic Women3.854.266.97
Another Race Women2.080.923.29
White Women13.135.0134.86
All Women29.7713.5351.53
Observations276,2101,667,763 
Note: Computations are for federal defendants sentenced under the U.S. Sentencing Guidelines in fiscal years 1994–2019. U.S. adult population averages computed over the years 1994–2019.
97 S. Cal. L. Rev. 299

Download

* Associate Professor of Law, Northwestern Pritzker School of Law. I am grateful to Joshua Braver, Samuel Buell, Franciska Coleman, Brandon Garrett, Michael Gentithes, Ben Grunwald, Andrew Hammond, Paul Heaton, Carissa Hessick, Christine Jolls, Kay Levine, James Lindgren, Yair Listokin, Yaron Nili, Lauren Ouziel, Maria Ponomarenko, John Rappaport, Megan Stevenson, Neel Sukhatme, Kegon Teng Kok Tan, Nina Varsava, Lisa Washington, Ron Wright, as well as participants at the 2022 CrimFest Conference, the 2022 Chicagoland Junior Scholars Conference, the 2022 Empirical Criminal Law Roundtable, the 2023 Annual Meeting of the American Law and Economics Association, the 2023 Conference on Empirical Legal Studies, the 2023 Harvard/Stanford/Yale Junior Faculty Forum, the Larry E. Ribstein Law & Economics Workshop at George Mason University Antonin Scalia Law School, the Soshnick Colloquium on Law and Economics at Northwestern Pritzker School of Law, and the University of Wisconsin-Madison La Follette School of Public Affairs Seminar for thoughtful comments on this work. Thomas Gordon and Matthew Marcin provided excellent research assistance. Finally, I thank the fantastic student editors of the Southern California Law Review for their meticulous and insightful editorial assistance.

Perfecting the Judicial Peremptory Challenge: A New Approach Using Preliminary Data on California Judges in 2021

Even the most carefully planned and genius strategies are pointless without an assumption of fairness: chess depends on a fair arbiter, soccer depends on a fair referee, and litigation depends on a fair judge. Just as arbiters and referees are frequently criticized for questionable decisions, judges also deal with accusations that bias has impermissibly clouded their judgment. To protect litigants, the California Legislature presented a solutionthe California Code of Civil Procedure section 170.6, a statute arming litigants with the option to replace their assigned judge if they declare that judge biased. This judicial peremptory challenge asks for no evidence of bias, further frustrating the disagreement between proponents who claim that this right will trigger a chain reaction to increased public confidence and decreased discrimination against litigants, and opponents who conversely warn that it will open a Pandora’s box of abuse, intimidation, and discrimination against innocent judges. The difficulty of constraining various harmful human tendencies is the problem of judicial peremptory challenges writ large.

It appears that much of this policy debate about judicial peremptory disqualification is informed by theory rather than empirical data. The study conducted by this Note reveals that, at least in 2021, (1) peremptory challenges do not occur often but abuse still occurs among the few times they are asserted, and (2) timing and form rules are weak procedural obstacles. My proposal acknowledges that judges are sometimes not the epitome of neutrality but takes issue with litigants who may inflict damage on undeserving judges and the adjudication generally. Instead of the current “no-questions-asked” regime, the recommended procedure is the following: after litigants receive judicial analytics, they can file the disqualification motion with an independent judge who will review both the motion and the challenged judge’s evidentiary explanation for factual and legal sufficiency. Admittedly, like its federal counterpart, this is not peremptory per se, but it is preferrable as it will perfect the peremptory challenge and diminish the risk of abuse even more than the current model.

INTRODUCTION

There is in each of us a stream of tendency, whether you choose to call it philosophy or not, which gives coherence and direction to thought and action. Judges cannot escape that current any more than other mortals.

—Justice Cardozo1Benjamin N. Cardozo, The Nature of the Judicial Process 12 (1964) (footnote omitted).

The Lady Justice sculptures that adorn the United States Supreme Court building serve as a reminder of the high standards to which we hold judges: her blindfold and scales represent unwavering impartiality.2Figures of Justice, Sup. Ct., https://www.supremecourt.gov/about/figuresofjustice.pdf [https://perma.cc/FPG3-33M4]. But juxtaposing Lady Justice, a godlike figure from ancient mythology,3Id. with judges, human beings vulnerable to inevitable fallibility,4Understandably, judges may find it challenging to be “patient, dignified and courteous” at all times given stressors in their personal life. Debra C. Weiss, Judge Agrees to Reprimand after Outbursts Directed at Plaintiff’s Attorney, Scheduling Clerk, ABA Journal (Sept. 26, 2022, 9:55 AM), https://www.abajournal.com/news/article/judge-agrees-to-reprimand-after-outbursts-directed-at-plaintiffs-attorney-scheduling-clerk [https://perma.cc/LZS9-K3D9]. For example, a magistrate judge in South Carolina self-reported himself to the Office of Disciplinary Counsel for using profanity in a comment directed at the plaintiff’s attorney and subsequently completed anger management counseling under the direction of the South Carolina Supreme Court. Id. At the time of his outburst, the judge was struggling to take care of his severely autistic son with epilepsy and his wife who had recently experienced serious health issues. Id. begs the question of whether these standards are unattainable ideals. What happens when judges cannot wear the blindfold and hold the scales yet still wield the sword symbolizing power?5Figures of Justice, supra note 2. The California Legislature responded to this reality by enacting California Code of Civil Procedure section 170.6 (“Section 170.6”) which grants judicial peremptory6“Peremptory” is defined as “putting an end to or precluding a right of action, debate, or delay” and “not providing an opportunity to show cause why one should not comply.” Peremptory, Merriam-Webster, https://www.merriam-webster.com/dictionary/peremptory [https://perma.cc/CKA8-53S2]; see also Peremptory, Legal Info. Inst., https://www.law.cornell.edu/wex/peremptory [https://perma.cc/7PLM-B7SZ] (“Peremptory means final and absolute, without needing any underlying justification.”). The alternative definition that “peremptory” means “expressive of urgency or command” seems befitting as well considering the nature of these challenges. Peremptory, Merriam-Webster, https://www.merriam-webster.com/dictionary/peremptory [https://perma.cc/CKA8-53S2]. challenges, or the power to automatically disqualify a judge for bias even without any evidence of such bias, to litigants.7Cal. Civ. Proc. Code § 170.6 (Deering 2023). Given that plaintiffs and defendants in the United States bear the burden of proof to succeed in their claims and defenses respectively, the significance of this exceptional legal right is apparent. But the California Legislature was not blind to the potential for this statute to act as a double-edged sword:8Johnson v. Superior Ct., 329 P.2d 5, 8 (Cal. 1958) (“The possibility that [Section 170.6] may be abused by parties seeking to delay trial or to obtain a favorable judge was a matter to be balanced by the Legislature against the desirability of the objective of the statute.”). litigants and their attorneys are naturally inclined to exploit this power to “shop” for a judge that is likely to favor their cause.9Consider former President Trump’s lawsuit against Hillary Clinton, among others, in which “Trump’s legal team . . . was specifically seeking out a particular federal judge: one he appointed as president.” Jose Pagliery, Trump Went Judge Shopping and It Paid Off in Mar-a-Lago Case, Daily Beast (Sept. 6, 2022, 11:07 AM), https://www.thedailybeast.com/donald-trump-went-judge-shopping-and-it-paid-off-in-mar-a-lago-case [https://perma.cc/VY8M-JMMK]. This cost-benefit analysis (“judge shopping,” which seems contradictory to the very essence of judging, weighed against public confidence in the judiciary) still plagues practitioners, legal academics, and judges today, decades after Section 170.6 was added to the California Code of Civil Procedure.

Although peremptory challenges are more commonly associated with jurors rather than judges,10See Peremptory Challenge, Legal Info. Inst., https://www.law.cornell.edu/wex/peremptory_challenge [https://perma.cc/XY5T-693M] (defining “peremptory challenge” only in the context of juror exclusion). the ability to change the assigned judge cannot be understated. The jury has indisputable influence over a case’s outcome by “mak[ing] findings of fact and render[ing] a verdict for [] trial.”11Jury, Legal Info. Inst., https://www.law.cornell.edu/wex/jury [https://perma.cc/7APK-JMEP]. Indeed, the foundational right to a judgment by one’s peers in the community dates back to the Magna Carta.12What Does the Magna Carta Mean?, Magna Carta, https://ipamagnacarta.org.au/what-does-magna-carta-mean [https://perma.cc/V6F9-F7SK]. Nonetheless, the judge still decides questions of law13Jury, supra note 11. and thus arguably holds equal, if not more, influence than the jury.14How Courts Work, Am. Bar Ass’n (Sept. 9, 2019), https://www.americanbar.org/groups/public_education/resources/law_related_education_network/how_courts_work/jury_role [https://perma.cc/2CC8-F4FC]. This is especially so given all cases have a judge but not all of them have a jury.15Bridey Heing, What Does a Juror Do? 7 (2018). In a bench trial without a jury, the judge “decides the facts of the case and applies the law.” Bench Trial, Legal Info. Inst., https://www.law.cornell.edu/wex/bench_trial [https://perma.cc/E4P7-BX5Q]. Unlike criminal cases in which defendants are guaranteed the right to a trial by jury under the Sixth Amendment of the U.S. Constitution, civil cases are not always afforded the same right. Jury, supra note 11. Moreover, a majority of cases do not proceed to trial: a judge’s ruling on a summary judgment motion has a conclusory effect akin to the end of trial.16Summary Judgment, Legal Info. Inst., https://www.law.cornell.edu/wex/summary_judgment [https://perma.cc/J8WW-HES5]. Even if a case reaches trial, a successful motion for judgment as a matter of law17Rule 50. Judgment as a Matter of Law in a Jury Trial; Related Motion for a New Trial; Conditional Ruling, Legal Info. Inst., https://www.law.cornell.edu/rules/frcp/rule_50 [https://perma.cc/7822-FHUR]. or a motion for new trial18Motion for New Trial, Legal Info. Inst., https://www.law.cornell.edu/wex/motion_for_new_trial [https://perma.cc/XJ9E-DBKR]. can subvert the jury’s verdict.

Considering judges’ unparalleled authority over litigants’ fate, notwithstanding the jury’s role, it is no surprise that judges must not “manifest bias . . . including but not limited to bias . . . based upon race, sex, gender, religion, national origin, ethnicity, disability, age, sexual orientation, marital status, socioeconomic status, or political affiliation . . . .”19Model Code of Jud. Conduct r. 2.3 (Am. Bar Ass’n 2020); see also Model Code of Jud. Conduct Canon 2 (Am. Bar Ass’n 2020) (“A judge shall perform the duties of judicial office impartially, competently, and diligently.”). Judicial independence not only has a rich history predating Enlightenment philosophy,20See, e.g., Ben W. Palmer, Books for Lawyers, 36 Am. Bar Ass’n J. 744, 768–69 (1950) (reviewing The Code of Maimonides: The Book of Judges (A.M. Hershman trans., 1949) to reveal how early Jewish law valued “perfect impartiality” in judges). but is also at the core of national identity in the United States: former President Adams, one of the Founding Fathers, wrote about the right to trial by “judges as free, impartial, and independent as the lot of humanity will admit” in the original Massachusetts Constitution.21Roy A. Schotland, New Challenges to States’ Judicial Selection, 95 Geo. L.J. 1077, 1079 (2007) (quoting John Adams in the original Massachusetts Constitution of 1780). Judges are supposed to represent the best of human nature, maintaining superior morals and ethics. This image erodes when judges rule with regard to “which side is popular” and “who is ‘favored.’ ”22How Courts Work, supra note 14. Impartiality in the courts is not a mere exercise in political correctness but a vital component of a fair, just, and democratic society rid of corruption. Once the public no longer trusts judges to treat them equally with their adversary, a domino effect to anarchy may ensue whereby people will stop respecting and therefore complying with orders from the judiciary and government at large. However, judicial discretion is as crucial to the proper functioning of the legal system as judicial impartiality because indeterminate laws require judges to “consider practical consequences and the overall context of a matter.”23David F. Levi, What Does Fair and Impartial Judiciary Mean and Why Is It Important?, Duke L. Bolch Jud. Inst. (Nov. 5, 2019), https://judicialstudies.duke.edu/2019/11/what-does-fair-and-impartial-judiciary-mean-and-why-is-it-important [https://perma.cc/RY6S-4NPW]. Alexander Hamilton, one of the Framers of the U.S. Constitution, distinguished between the “guided exercise of discretion” and the “imposition of personal will and preference” by highlighting the “importance of courageous judges to the preservation of individual liberty and to the amelioration of oppressive legislation.”24Id.

Ideally, litigants would always use Section 170.6 in good faith to defend themselves from judicial bias. Unfortunately, courts confront the ironic truth that some litigants abuse this ability as an offensive maneuver instead. Litigants may take advantage of peremptory challenges to substitute their judge with one that has aligned interests—that is, a biased judge. Section 170.6 can accordingly exacerbate the very problem it was designed to minimize. Bias is a two-way street in which litigants can also discriminate against judges of a particular gender, race, or ethnicity, among other demographics. There was increased legislative movement toward eliminating peremptory juror challenges for this reason in 202125See, e.g., S. 212, 2021 Leg., Reg. Sess. (Cal. 2021); S. 2211, 2021 Leg., Reg. Sess. (Miss. 2021); S. S6066, 2021 Leg., Reg. Sess. (N.Y. 2021). and publicity on race-based discrimination in jury selection in 2022.26See, e.g., Janet Miranda, Race-Based Jury Strikes at Issue in New Texas Supreme Court Case, Bloomberg L. (Sept. 2, 2022, 11:31 AM), https://www.bloomberglaw.com/bloomberglawnews/us-law-week/XHUTRFG000000 [https://perma.cc/H5HR-SRN4] (reporting on a controversial case in which attorneys peremptorily challenged all of the white, male jurors); Jason Meisner & Megan Crepeau, Jury in R. Kelly’s Chicago Federal Case Selected; Opening Statements Set for Wednesday, Chi. Trib. (Aug. 16, 2022, 7:43 PM), https://www.chicagotribune.com/news/criminal-justice/ct-r-kelly-chicago-federal-trial-jury-selection-day-two-20220816-i2gavfvjm5cp5enqy2cwulzpbq-story.html [https://perma.cc/8CNS-YGBG] (“Things got testy when Kelly’s lead attorney . . . successfully challenged three of the prosecution’s strikes of Black jurors, alleging they were based solely on race.”). If attorneys can discriminate against potential jury members, they can discriminate against judges as well, and the legal field should brace for any future spillover on peremptory challenges to judges. In 2022, 60.1% of judges in California were male, and 61.4% of them were white.27Jud. Couns. of Cal., Demographic Data Provided by Justices and Judges 1 (2022), https://www.courts.ca.gov/documents/2023-JO-Demographic-Data.pdf [https://perma.cc/Q6CU-UU6P]. Imagine the harm that would result if most of the disqualified judges were members of groups that have historically endured discrimination. The judiciary would subsequently lose the diversity of thought and experiences necessary to adequately understand and evaluate heterogeneous litigants from the United States, a country often referred to as a melting pot.

This Note illustrates the need to abandon the judicial peremptory challenge as it exists today and instead opt for a blend of other variations—specifically, the challenge should be less peremptory and more stringent. Preliminary empirical data in 2021 reveals that (1) peremptory challenges do not occur frequently but abuse still occurs among the few times they are asserted, and (2) timing and form rules are weak procedural obstacles. Although the challenge is not widely abused, a different model will decrease the incidences of abuse even further. In lieu of a conclusory allegation of bias that is instantaneously granted, the proposed disqualification approach allows the challenged judge to refute the allegation with evidentiary explanations. This will hopefully pull the reins on the litigants, however few, who make an unwarranted, illusory charge of bias against their judge in order to gain a tactical advantage.

This Note begins by providing a high-level overview of how peremptory challenges to judges are treated by federal courts and other state courts besides California. It also explores Section 170.6 in detail, particularly the statute’s legislative history and interplay with judicial rules and peremptory juror challenges. Next, it summarizes the current policy arguments both in favor and against peremptory disqualification of judges: points of contention include discrimination against judges and confidence in the judiciary, among others. It continues with an analysis of data collected from every order in 2021 in which a California superior court judge decided on a Section 170.6 motion, tracking for the number of filed motions, number of denied motions and why they were rejected, number of disqualified judges, and the disqualified judges’ political party. It then synthesizes the findings with judicial disciplinary actions due to bias in 2021, which informs the policy debate by revealing the concerns that actually come to fruition in practice, rather than in theory only, at least in the context of California for this time frame. It additionally explores the reasons behind challenging a judge using The Robing Room, a public forum. Afterwards, it discusses alternative disqualification procedures offered by some legal scholars before advocating a new approach. Finally, the Note ends with recommendations for future research.

I.  MODERN LAW OF JUDICIAL PEREMPTORY DISQUALIFICATION

Section 170.6 is a relatively recent addition to judicial disqualification law28Act of 1957, ch. 1055, 1957 Cal. Stat. 2288, https://clerk.assembly.ca.gov/sites/clerk.assembly.ca.gov/files/archive/Statutes/1957/57Vol1_57Chapters.pdf#page=2 [https://perma.cc/3MFS-HRVV].—the decades since its enactment pale in comparison to the more than one thousand years people have spent developing legal justifications for disqualifying judges.29See, e.g., The Codex of Justinian 619 (Bruce W. Frier & Serena Connolly, eds., Fred H. Blume trans., 2016) (stating that Roman law allowed for judicial disqualification if it occurred before trial). In the mid-eighteenth century, the thirteen American colonies adopted English jurisprudence that,30John P. Frank, Disqualification of Judges, 56 Yale L.J. 605, 609 (1947). unlike civil law countries, narrowed the scope of judicial disqualification so that a judge could only face disqualification if they had a direct pecuniary interest in the case.31Richard E. Flamm, Judicial Disqualification: Recusal and Disqualification of Judges 6 (2d ed. 2007). Thus, lacking basis in common law,32Frank, supra note 30, at 612. disqualification for bias did not enter the stage until 1903, well after the founding of the United States, when Montana’s legislature answered the cries of a losing litigant.33See id. at 608 n.8. This win for victims of judicial bias was part of a growing focus on ensuring that judges apply the law in an evenhanded manner,34See, e.g., Act of Mar. 3, 1821, ch. 51, 3 Stat. 643 (ordering recusal if a judge believes they are so related or connected to a party that their decision would be improper) (codified at 28 U.S.C. § 144); Act of Mar. 3, 1891, ch. 517, § 3, 26 Stat. 826, 827 (forbidding a judge from hearing the appeal of a case they tried) (codified at 28 U.S.C. § 47). eventually escalating into the federal law’s official acknowledgment.35Act of Mar. 3, 1911, ch. 231, § 21, 36 Stat. 1087, 1090 (allowing disqualification if a party files a sufficient affidavit asserting bias) (codified at 28 U.S.C. § 144). The evolution of judicial disqualification finds itself at a fork in the road: some states in the West and Midwest, including California, allow disqualification with an allegation of bias alone—known as a peremptory challenge—while other states in the East and South join the federal courts in imposing stricter standards by requiring support for the allegation as well.36See, e.g., Cal. Civ. Proc. Code § 170.6 (Deering 2023); 725 Ill. Comp. Stat. Ann. 5/114–5 (LexisNexis 2023); N.Y. Jud. Law § 14 (Consol. 2023); Tex. Gov’t Code Ann. § 25.00255 (LexisNexis 2023); Wamser v. State, 587 P.2d 232, 234–35 (Alaska 1978) (“In the absence of a challenge for cause, no such right [to peremptory challenges] existed at common law, and it is not afforded in the federal courts or in many states in the absence of a showing of factual bias.” (footnotes omitted)).

A.  Federal Law

In 1911, 28 U.S.C. § 144 introduced judicial peremptory challenges into the federal realm.37See, e.g., Alan J. Chaset, Disqualification of Federal Judges by Peremptory Challenge 5–6 (1981) (“[28 U.S.C. § 144] has remained virtually unchanged since it was enacted in 1911.” (footnote omitted)). This federal statute closely mirrors Section 170.6 as it permits the disqualification of a district court judge upon a timely affidavit claiming bias. However, it departs from Section 170.6 in a significant way: it requires the affidavit to “state the facts and the reasons for the belief that bias or prejudice exists” and accordingly affords less leeway to litigants.3828 U.S.C. § 144. On its face, its wording and legislative history hint at the intent for peremptory disqualification;39Chaset, supra note 37, at 7 n.11 (“Congressman Cullop of Indiana, the chief sponsor of the legislation, [stated that 28 U.S.C. § 144] ‘provides that the [challenged] judge shall proceed no further with the case.’ ” (citing 46 Cong. Rec. 2627 (1911)); Charles Gardner Geyh & Kris Markarian, Judicial Disqualification 83 (2010) (“Such an interpretation would render [28 U.S.C. § 144] akin to peremptory disqualification procedures . . . and the legislative history of [28 U.S.C. § 144] lends some support for this interpretation.”); Debra Lyn Basssett, Judicial Disqualification in the Federal Appellate Courts, 87 Iowa L. Rev. 1213, 1224 n.54 (2002) (“Congress modeled the federal statute on an Indiana statute, which provided for automatic disqualification upon the filing of the affidavit.”). judicial interpretation steered it on the opposite trajectory.40Frank, supra note 30, at 629 (“Frequent escape from the statute has been effected through narrow construction of the phrase ‘bias and prejudice.’ ”). Judges are incentivized to narrowly interpret the statute when applying it to themselves. Amanda Frost, Keeping Up Appearances: A Process-Oriented Approach to Judicial Recusal, 53 U. Kan. L. Rev. 531, 551 (2005). One attorney argued that 28 U.S.C. § 144 should be amended to include a “clear directive that the federal peremptory disqualification statute is to be construed liberally in favor of disqualification, and not as a nit to be picked until the peremptory purpose of the statute is eviscerated by judicial interpretation”; otherwise, it should be repealed so the other federal judicial disqualification statute, 28 U.S.C. § 455, can take the lead. Richard E. Flamm, History of and Problems with the Federal Judicial Disqualification Framework, 58 Drake L. Rev. 751, 763 (2010). After the Supreme Court in Berger v. United States opined that the challenged judge may conduct a hearing to scrutinize the alleged facts for legal sufficiency,41Berger v. United States, 255 U.S. 22, 32 (1921). the Court clarified in Liteky v. United States that “expressions of impatience, dissatisfaction, annoyance, and even anger, that are within the bounds of what imperfect men and women, even after having been confirmed as federal judges, sometimes display” fall short of bias.42Liteky v. United States, 510 U.S. 540, 555–56 (1994). The latter case defined the extrajudicial source doctrine: critical, disapproving, or hostile opinions based on facts or events during the proceedings do not warrant disqualification unless “they reveal such a high degree of favoritism or antagonism as to make fair judgment impossible.”43Id. at 555. It is worth mentioning that the Ninth Circuit also adds a reasonable person test. Pesnell v. Arsenault, 543 F.3d 1038, 1043 (9th Cir. 2008). Congress largely acquiesced to this rejection of peremptory intent lest they infringe upon the separation of powers by attempting to regulate the judiciary.44Flamm, supra note 40, at 756 (“Congress could have taken steps to disabuse the federal judiciary of this notion, but it did not.”); Frost, supra note 40, at 551–52 (“The legislative and executive branches may feel that it is inappropriate to dictate the minutiae of procedures to be followed when litigants seek to remove a judge from a case, preferring to leave it to the judiciary to clean its own house.”). As a result, “disqualification under [28 U.S.C. § 144] has been rare.”45Gabriel D. Serbulea, Due Process and Judicial Disqualification: The Need for Reform, 38 Pepp. L. Rev. 1109, 1125 (2011); see Geyh & Markarian, supra note 39, at 83. Naturally, the statute could no longer be classified as fully peremptory, distinguishing it from its state counterparts that order automatic disqualification, like Section 170.6.

B.  California Law

1.  California Code of Civil Procedure Section 170.6

In 1957, the California legislature debated whether to accept or deny the legacy of judicial peremptory challenges and ultimately concluded with the birth of Section 170.6 through an “overwhelming vote of both houses of the Legislature” and approval by the Governor.46Johnson v. Superior Ct., 329 P.2d 5, 7 (Cal. 1958); Act of 1957, ch. 1055, 1957 Cal. Stat. 2288, https://clerk.assembly.ca.gov/sites/clerk.assembly.ca.gov/files/archive/Statutes/1957/57Vol1_57Chapters.pdf#page=2 [https://perma.cc/3MFS-HRVV]. The legislation was more radical47See, e.g., California Judges Benchbook: Civil Proceedings-Before Trial § 7.2 (West 2022) (“The right to exercise a peremptory challenge against a judge is a creation of statute: it did not exist before the enactment of [Section 170.6].”). than California Code of Civil Procedure section 170.1 which concerns challenges for cause48Cal. Civ. Proc. Code § 170.1 (Deering 2023); CCP 170.6 – Disqualification of a Judge on Grounds of Prejudice, Shouse Cal. L. Grp., https://www.shouselaw.com/ca/defense/disqualification-of-judge-for-prejudice [https://perma.cc/LC5B-E6W8] (“Under [California Code of Civil Procedure section 170.1], a judge can be removed ‘for cause’ if any one or more of the following are true: the judge has personal knowledge of disputed facts in the case, the judge served as an attorney in the proceeding or advised a party in the proceeding, the judge has a financial interest in the proceeding, the judge, or the judge’s spouse, is a party in the case or an officer, director, or trustee of a party, or the judge, or a person related to the judge, is associated in private practice of law with an attorney in the case.”). Section 170.1 also permits self-removal if the judge believes their recusal would “further the interests of justice” or their impartiality is at risk. Id. Unlike Section 170.6, there are no limits on the number of challenges, Disqualification of a Judge for Prejudice, Eisner Gorin LLP, https://www.egattorneys.com/disqualification-of-a-judge [https://perma.cc/M72S-TTU7], and specific proof is required, How to Request to Change Your Judge, Res. Ctr for Self-Represented Litigants, https://www.courts.ca.gov/partners/documents/request_change_judge.doc [https://perma.cc/F764-FBN8]. See generally O’Connor’s California Practice Civil Pretrial Ch. 2-D § 3 (West 2023). and California Code of Civil Procedure section 170.5 (added as section 170.4 in 1897), which addresses bias as a ground for disqualification.49Civ. Proc. § 170.5 (Deering 2023); Johnson, 329 P.2d at 7–8. This was not the first time the Legislature dealt with judicial peremptory challenges: four previous measures failed to receive executive approval despite passage by the Legislature.50The four measures are A.B. 442 passed in 1941, A.B. 479 passed in 1951, S.B. 392 passed in 1953, and S.B. 89 passed in 1955. Johnson, 329 P.2d at 7 n.2. Therefore, Johnson v. Superior Court, the first case to apply Section 170.6, acknowledged how the “[s]tate [b]ar and the Legislature have long felt that there is a need for such a measure.”51Johnson, 329 P.2d at 7.

Unlike federal judges under 28 U.S.C. § 144, California judges generally fortified Section 170.6 by “liberally constru[ing it] with a view to effect its objects and to promote justice,”52Le Louis v. Superior Ct., 257 Cal. Rptr. 458, 466 (Ct. App. 1989); see, e.g., Pappa v. Superior Ct., 353 P.2d 311, 314–15 (Cal. 1960) (“[L]imiting each ‘side’ to one challenge [of a judge for prejudice] . . . does not arbitrarily discriminate against multiple parties,” since “[t]he Legislature could reasonably determine that this limited restriction was justified in order to prevent undue delays which could otherwise occur.” (citing Johnson, 329 P.2d at 5)); Mayr v. Superior Ct., 39 Cal. Rptr. 240, 243 (Ct. App. 1964) (“[Section 170.6] should not be so strictly construed that the legislative will is thwarted.”); Solberg v. Superior Ct., 561 P.2d 1148, 1159 (Cal. 1977) (“[Section 170.6] makes no provision for a detailed statement of facts, and it is reasonable to infer the Legislature did not intend to impose such a condition.”). starting with its constitutionality. Since the constitutionality of peremptorily disqualifying a judge has been debated since the early twentieth century,53Annotation, Constitutionality of Statute Making Mere Filing of Affidavit of Bias or Prejudice Sufficient to Disqualify Judge, 5 A.L.R. 1275 (1920) (summarizing cases that declared peremptory challenges of judges either constitutional or unconstitutional). it comes as no surprise that Section 170.6 came under attack almost immediately after its enactment. Even before the Legislature took action, the courts in the state ruled in several cases that a similar disqualification statute enacted in 1937 was unconstitutional.54Annotation, Constitutionality of Statute Which Disqualifies Judge upon Peremptory Challenge, 115 A.L.R. 855 (1938) (discussing how Austin v. Lambert, 77 P.2d 849 (Cal. 1938), Daigh v. Schaffer, 73 P.2d 927 (Cal. 1937), and Krug v. Superior Ct., 77 P.2d 854 (Cal. 1938), determined that the older disqualification statute from 1937 was unconstitutional). Johnson represented a turning point as the Supreme Court of California deemed Section 170.6 constitutional and overruled the lower court’s decision that “the statute makes an unconstitutional delegation of legislative and judicial powers to litigants and their attorneys and is an unwarranted interference with the powers of the courts.”55Johnson, 329 P.2d at 7. Section 170.6 is materially different from its unconstitutional predecessor because it calls for litigants to submit a sworn statement instead of a “judicial determination of the existence of the fact.”56Id. at 8–9 (“[The disqualification statute enacted in 1937] provided for a ‘peremptory challenge’ of the judge assigned to hear the case without requiring the person making the challenge to state the ground for his objection or to make a declaration under oath that the ground in fact existed.”). According to the court, Section 170.6 complies with the Constitution and deserves protection because “[p]rejudice, being a state of mind, is very difficult to prove, and, when a judge asserts that he is unbiased, courts are naturally reluctant to determine that he is prejudiced.”57Id. at 8. About twenty years later, Section 170.6’s constitutionality returned to the forefront in Solberg v. Superior Court—this court found no separation of powers violation under California Constitution Article III, Section 3.58Solberg v. Superior Ct., 561 P.2d 1148, 1162 (Cal. 1977). In a post-Johnson and Solberg world, the conversation between Section 170.6’s proponents and opponents has shifted away from constitutionality, but policy concerns persist. As this Note will later discuss, the thousand-year-old debate has still not found its rest.

Section 170.6 was amended to widen its scope: beginning in 1959, the statute extended to criminal, not just civil, cases,59Act effective Sept. 18, 1959, ch. 640, 1959 Cal. Stat. 2620, 2620. This amendment settled the dispute regarding whether withholding this right from criminal parties was unconstitutional discrimination under the Fourteenth Amendment of the U.S. Constitution and the California Constitution under Article I, Sections 11 and 21, and Article IV, Section 25 for unreasonable classifications. See Johnson, 329 P.2d at 9. and beginning in 1961, oral statements under oath, not just written documents.60Act effective Sept. 15, 1961, ch. 526, sec. 1, § 170.6(2), 1961 Cal. Stat. 1628, 1629. This trend halted in 1965, when the Legislature forbade litigants from receiving a judicial reassignment if their original judge already presided over a proceeding prior to trial that involved a “determination of contested fact issues relating to the merits.”61Act of 1965, ch. 1442, sec. 1, § 170.6(2), 1965 Cal. Stat. 3375, 3375–76; see Bambula v. Superior Ct., 220 Cal. Rptr. 223, 224 (Ct. App. 1985) (“This addition preserves the right of a party to disqualify a judge under [the statute,] notwithstanding the fact that the judge had heard and determined an earlier demurrer or motion, or other matter not involving ‘contested fact issues’ relating ‘to the merits’ without challenge in the same cause.”). For the next ten or so years, the statute was only amended twice—in 196762Act of 1967, ch. 1602, sec. 2, § 170.6(1), 1967 Cal. Stat. 3832, 3832. It also added the option of including a “declaration under penalty of perjury.” Id. at sec. 2, § 170.6(2) at 3833. and 197663Act of 1976, ch. 1071, sec. 1, § 170.6(1), 1976 Cal. Stat. 4814, 4815.—to subject court commissioners and referees to potential peremptory disqualification as well.64Although Section 170.6 applies to judges, court commissioners, and referees of a superior, municipal, or justice court, it does not affect a superior court judge who is appointed by an appellate court as a referee. People v. Gonzalez, 800 P.2d 1159, 1197 n.44 (Cal. 1990). The Legislature obviously did not shy away from its peremptory intent, given that the affidavit form was amended in 1981 to add “peremptory challenge.”65Act of 1981, ch. 192, sec. 1, § 170.6(5), 1981 Cal. Stat. 1116, 1117–18. After another amendment in 1982 that clarified the timeliness requirement for single-judge systems,66Act of 1982, ch. 1644, sec. 2, § 170.6(2), 1982 Cal. Stat. 6678, 6682–83. the statute was expanded yet again in 1985. Now, litigants who file an appeal that results in the reversal of the trial court’s judgment qualify for protection if the “trial judge in the prior proceeding is assigned to conduct a new trial on the matter.”67Act of 1985, ch. 715, sec. 1, § 170.6(2), 1985 Cal. Stats. 2350, 2351. Timeliness was then defined as ten days for criminal cases with an all-purpose assignment in 1989.68Act of 1989, ch. 537, sec. 1, § 170.6(2), 1989 Cal. Stats. 1803, 1803–04. The following two amendments in 199869Act of 1998, ch. 167, sec. 1, § 170.6(1), 1998 Cal. Stats. 932, 932–33. and 200270Act of 2002, ch. 784, sec. 36, § 170.6(1), 2002 Cal. Stats. 4710, 4744. There was also the technical change of updating the year on the affidavit form. Id. at sec. 36, § 170.6(5) at 4746–47. responded to modifications of the California Constitution—the elimination of the justice court71Cal. Const. art. VI, §§ 1, 5(b) (§ 5 repealed 2002). and unification of the municipal and superior courts, respectively72Cal. Const. art. VI, § 5(3) (repealed 2002).—which were products of the Legislature’s “stead[y] move[ment] towards completion of the courts’ restructuring.”73Senate Judiciary Comm., SB 1316 Senate Floor Analyses, at 2 (Cal. 2002). In 2003, the Legislature merely maintained the codes74Senate Judiciary Comm., SB 600 Senate Floor Analyses, at 2 (Cal. 2003) (“Each year, the Legislative Counsel’s Office identifies grammatical errors and other errors of a technical nature that have been inadvertently enacted into statutory law.”). and did not make any substantive changes.75Act of 2003, ch. 62, sec. 22, § 170.6, 2003 Cal. Stats. 264. The last amendment, in 2010, made similar corrections, but also extended the filing deadline for civil cases with an all-purpose assignment to fifteen days after receiving notice of the assignment76State Assembly 1894, 2010 Leg., Reg. Sess. (Cal. 2010). There was a need to reconcile the Code of Civil Procedure and the Trial Court Delay Reduction Act of 1990. Cal. Assembly Judiciary Comm., AB 1894 Assembly Floor Analysis, at 2 (Cal. 2010). and “codif[ied] existing court practices by requiring the party making the challenge to notify all other parties within five days after making the motion [to peremptorily disqualify the judge].”77Cal. Assembly Judiciary Comm., AB 1894 Assembly Floor Analysis, at 2 (Cal. 2010).

In general, Section 170.6 guarantees litigants the extraordinary right to have an alternate superior court judge hear their matter once they accuse their judge78This covers both retired judges who are assigned to temporarily act as a regular sitting judge to hear a case and active, full-time judges. People v. Superior Ct. (Mudge), 62 Cal. Rptr. 2d 721, 725 (Ct. App. 1997). of bias, even without any factual basis for actual bias.79General legal conclusions will do. Andrews v. Joint Clerks Port Lab. Rels. Comm., 48 Cal. Rptr. 646, 651 (Ct. App. 1966); People v. Rodgers, 121 Cal. Rptr. 346, 347 (Ct. App. 1975); CCP § 170.6 – Disqualification of a Judge on Grounds of Prejudice, supra note 48. See generally O’Connor’s California Practice Civil Pretrial, supra note 48, at Ch. 2-D § 4. Litigants can raise a challenge under Section 170.6 at any trial, special proceeding, or hearing involving a “contested issue of law or fact,”80Cal. Civ. Proc. Code § 170.6(a)(1) (Deering 2023); Andrews, 48 Cal. Rptr. at 650–51; Est. of Cuneo, 29 Cal. Rptr. 497, 499 (Ct. App. 1963). From a policy standpoint, this stops litigants from seeking more favorable rulings from a different judge. People v. Richard, 149 Cal. Rptr. 344, 347 (Ct. App. 1978); People v. Paramount Citrus Ass’n, 2 Cal. Rptr. 216, 221 (Ct. App. 1960); Dennis v. Overholtzer, 3 Cal. Rptr. 458, 459 (Ct. App. 1960). including trial, law and motion proceedings, injunction hearings, and contested probate or family law proceedings, but excluding settlement or case management conferences.81Peremptory Challenge of a Judge: Remove the Judge from Your Case, Sacramento Cnty. Pub. L. Libr. 1 (Nov. 2021), https://saclaw.org/wp-content/uploads/sbs-peremptory-challenge-of-a-judge.pdf [https://perma.cc/GT8J-FEAR]. Litigants should not disregard local county rules—special courts like Dependency Court and Family Court might restrict or completely forbid peremptory challenges in certain types of proceedings. Id. Each side in a case, defined by whether the co-plaintiffs or co-defendants have substantially adverse interests,82Pappa v. Superior Ct., 353 P.2d 311, 314 (Cal. 1960) (“The privilege conferred by section 170.6, unlike the right to counsel, may be exercised by more than one codefendant only where they have substantially adverse interests, and obviously the mere fact that they choose to be represented by separate counsel does not show that such a conflict of interests exists.”). If co-parties share interests, but one party already moved forward with a challenge without the other parties’ consent, they all lose their one challenge. Louisiana-Pacific Corp. v. Philo Lumber Co., 210 Cal. Rptr. 368, 369 (Ct. App. 1985). is given one challenge—the norm.83Note that challenges for cause through California Code of Civil Procedure section 170.1 are still available after exhausting the peremptory challenge. Serbulea, supra note 45, at 1144. “[I]f the trial judge in the prior proceeding is assigned to conduct a new trial84If the issue to be resolved on remand requires the court to perform “merely a ministerial act,” there is no “new trial” within the meaning of Section 170.6. Stegs Invs. v. Superior Ct., 284 Cal. Rptr. 495, 495 (Ct. App. 1991); Overton v. Superior Ct., 27 Cal. Rptr. 2d 274, 275 (Ct. App. 1994). The “new trial” does not have to take place after trial: it can occur after any kind of final judgment, such as summary judgment. Stubblefield Constr. Co. v. Superior Ct., 97 Cal. Rptr. 2d 121, 124 (Ct. App. 2000). on the matter” after “reversal on appeal of a trial court’s final judgment,” the movant can still use Section 170.6 regardless of whether they have already availed themselves of this procedure.85Civ. Proc. § 170.6(a)(2) (emphasis added). They, however, cannot make the motion “for the first time in post-trial matters which are essentially a ‘continuation’ of the main proceeding,”86Solberg v. Superior Ct., 561 P.2d 1148, 1158 (Cal. 1977). meaning “action[s] . . . involv[ing] ‘substantially the same issues’ and ‘matters necessarily relevant and material to the issues involved in the original action.’ ”87Matthews v. Superior Ct., 42 Cal. Rptr. 2d 521, 523 (Ct. App. 1995). A proceeding can be a continuation even if it has a different county clerk’s file number. Andrews v. Joint Clerks Port Lab. Rels. Comm., 48 Cal. Rptr. 646, 653–54 (Ct. App. 2012). The court in Pickett v. Superior Court, 138 Cal. Rptr. 3d 36, 42 (Ct. App. 2012), opined that the second plaintiff’s action was not a continuation of the first plaintiff’s action despite both actions alleging the same wrongful conduct because the second action sought additional relief. Likewise, in Bravo v. Superior Court, 57 Cal. Rptr. 3d 910, 914 (Ct. App. 2007), the instant case was not a continuation even though it concerned the same plaintiff and defendant because “the [second] action [arose] out of later events distinct from those in the previous action.” Absent good cause, there is no continuance of the trial or hearing because of the motion; if a continuance is granted for other reasons, the matter must be continued for limited periods to be reassigned as soon as possible.88Civ. Proc. § 170.6(a)(4). In the aftermath of the 2010 amendment, civil litigants must serve notice on all parties within five days of making the motion.89Id. § 170.6(a)(3).

Either an affidavit accompanied with a declaration that a “fair and impartial hearing or trial cannot take place” under penalty of perjury90Tyler Perez, Disqualifying a Judge: An Early Strategic Move, CMF (Mar. 30, 2023), https://cafamlaw.com/disqualifying-a-judge-an-early-strategic-move [https://perma.cc/SW8Q-MCM6]. or an oral motion under oath will suffice as long as it is made before the hearing or trial commences.91Civ. Proc. § 170.6(a)(2) (“In no event shall a judge, court commissioner, or referee entertain the motion if it is made after the drawing of the name of the first juror, or if there is no jury, after the making of an opening statement by counsel for plaintiff, or if there is no opening statement by counsel for plaintiff, then after swearing in the first witness or the giving of any evidence or after trial of the cause has otherwise commenced. If the motion is directed to a hearing, other than the trial of a cause, the motion shall be made not later than the commencement of the hearing.”); Haldane v. Haldane, 26 Cal. Rptr. 670, 675 (Ct. App. 1962). Litigants should submit the written or oral motion “as soon as possible after [they] know[] with some reasonable certainty who the actual trial judge will be”92Augustyn v. Superior Ct., 231 Cal. Rptr. 298, 302 (Ct. App. 1986); see Lawrence v. Superior Ct., 253 Cal. Rptr. 748, 751 (Ct. App. 1988) (“Knowledge of the assignment does not mean actual knowledge on the part of the party or his attorney but only that, upon further investigation or inquiry, the identity of the judge assigned to a particular department is ascertainable.”). and must take care to abide by the timing rules governing master calendar, all-purpose, and single-judge systems. Table 1 below describes each of these various case-management systems and their nuances:

Table 1.
Case-Management SystemDefinitionRules
Master CalendarThe case is assigned to different departments for specific types of matters.aThe litigant should make the motion to the judge supervising the master calendar. When the case is set for immediate trial, litigants are instructed to make their challenge at the time of assignment, but when the case is set for a later trial date, they should comply with the 10-day/5-day rule.b The 10-day/5-day rule dictates that a litigant must make their challenge at least five days before the trial or hearing date if the judge’s identity is known at least ten days before the trial or hearing date.c If they wait until they appear before the judge, they will have essentially waived their right to challenge that judge.d However, a late-appearing or late-named party will not be penalized as long as they make the motion within ten days after their appearance.e
All-Purpose/ Direct-CalendarThe randomly assigned judge “maintain[s] [their] own calendar, set[s] and handl[es] all motions and other proceedings, and conduct[s] trial.”f At the time of the assignment, the judge must be expected to process all substantial matters in addition to trial.gOnce a civil litigant receives noticeh of an all-purpose assignment, they have fifteen days to make their challenge.i If the litigant receives service by mail, they are entitled to a five-day extension.j However, if the litigant has not yet appeared, they have fifteen days after their appearance.k For a criminal litigant, they have ten days instead of fifteen days.l
Single-JudgeThere is only one judge in the courts.mA litigant must make their challenge within thirty days after they first appear in the action.n
Sources:  a  O’Connor’s California Practice Civil Pretrial, supra note 48, at Ch. 3-E § 3.3(2)(a). b  See, e.g., People v. Roerman, 10 Cal. Rptr. 870, 878–79 (Ct. App. 1961) (rejecting the motion because it was not made until the day trial was scheduled to begin even though the case was calendared to the judge for more than a month). c  E.g., Eagle Maint. & Supply Co. v. Superior Ct., 16 Cal. Rptr. 745, 747 (Ct. App. 1961). d  See, e.g., Michaels v. Superior Ct., 7 Cal. Rptr. 858, 860–61 (Ct. App. 1960); Peremptory Challenges to a Judge in California, L. Off. of Stimmel, Stimmel & Roeser, https://stimmel-law.com/en/articles/peremptory-challenges-judge-california [https://perma.cc.6NLU-5PLS]. e  Sch. Dist. of Okaloosa Cnty v. Superior Ct., 68 Cal. Rptr. 2d 612, 612 (Ct. App. 1997). f  O’Connor’s California Practice Civil Pretrial, supra note 48, at Ch. 3-E § 3.3(2)(b). g  See, e.g., People v. Superior Ct. (Lavi), 847 P.2d 1031, 1043 (Cal. 1993). h  According to the opinion of Cybermedia, Inc. v. Superior Ct., 82 Cal. Rptr. 2d 126, 127 (Ct. App. 1999), notice should reference the case name and full case number and be addressed to the attorney if the party is represented; otherwise, it is insufficient. i  Cal. Civ. Proc. Code § 170.6(a)(2) (Deering 2023). j  Motion Picture & Television Fund Hosp. v. Superior Ct., 105 Cal. Rptr. 2d 872, 876 (Ct. App. 2001). k  Civ. Proc. § 170.6(a)(2). l  Id. m  O’Connor’s California Practice Civil Pretrial, supra note 48, at Ch. 3-E § 3.3(2)(b). n  See, e.g., People v. Superior Ct. (Smith), 235 Cal. Rptr. 482, 484 (Ct. App. 1987) (refusing the movant’s argument that their motion was timely because it was made within thirty days after their attorney’s first appearance).

If litigants who successfully appeal the trial court’s judgment end up with the same trial judge, they must make the motion “within [sixty] days after the party or the party’s attorney has been notified of the assignment,”93Civ. Proc. § 170.6(a)(2). or else it will be time-barred. The motion must be directed to the very judge under attack (the particular department will not suffice) who will then determine if it has been duly presented.94See, e.g., Fry v. Superior Ct., 166 Cal. Rptr. 3d 328, 333 (Ct. App. 2013) (denying a peremptory challenge that was not made to anyone).

There are compelling policy reasons for these rules, namely that both litigants who “wish[] to postpone [the] motion until [they are] fully informed” and the court that needs “time to make adjustments after a disqualification” are satisfied.95See, e.g., L.A. Cnty. Dept. of Pub. Soc. Servs. v. Superior Ct., 138 Cal. Rptr. 43, 46 (Ct. App. 1977). Also, criminal litigants have less time to submit their motions than civil litigants because, like the Legislature probably thought, there are heightened concerns of abuse in criminal cases96Johnson v. Superior Ct., 329 P.2d 5, 9 (Cal. 1958). in which “the sides . . . are not in symmetrical positions,” as the prosecution possesses more power.97Anna Roberts, Defense Counsel’s Cross Purposes: Prior Conviction Impeachment of Prosecution Witnesses, 87 Brook. L. Rev. 1225, 1238 (2022). Criminal cases usually involve juries that dilute the judge’s influence, whereas civil cases are usually wholly decided by the judge.98The Differences Between a Criminal Case and a Civil Case, FindLaw, https://www.findlaw.com/criminal/criminal-law-basics/the-differences-between-a-criminal-case-and-a-civil-case.html [https://perma.cc/8NL2-5CKR]. Even so, criminal defendants are insulated from the “depriv[ation] of life, liberty, or property, without due process of law” by the Bill of Rights in the Fifth Amendment of the U.S. Constitution.99Unbiased Judge, Legal Info. Inst., https://www.law.cornell.edu/constitution-conan/amendment-5/unbiased-judge [https://perma.cc/DR8V-KLFE]. Given what they have to lose as compared to civil litigants, it is critical to avoid infringement on their right to a fair trial.

If the motion is timely filed with acceptable form,100Cal. Civ. Proc. Code § 170.6(a)(6) (Deering 2023) and Peremptory Challenge to Judicial Officer (Code Civ. Proc., § 170.6), Superior Ct. of Cal., Cnty. Of L.A., https://www.lacourt.org/forms/pdf/laciv015.pdf [https://perma.cc/3EBJ-9QME], provide a template for the motion. For examples of a Motion for Peremptory Challenge, Declaration in Support of Peremptory Challenge, and Order of Transfer, see Peremptory Challenge of a Judge: Remove the Judge from Your Case, supra note 81, at 5–10. it will be granted, and the transition process is automatic in the sense that the affidavit is not contestable. The “judge immediately loses jurisdiction over the case,” and “any action that [they] make[] in the case [is] considered ‘void,’ ” save to transfer the case to another judge.101CCP § 170.6 – Disqualification of a Judge on Grounds of Prejudice, supra note 48; see, e.g., Est. of Cuneo, 29 Cal. Rptr. 497, 499 (Ct. App. 1963); Woodman v. Selvage, 69 Cal. Rptr. 687, 691 (Ct. App. 1968); Andrews v. Joint Clerks Port Lab. Rels. Comm., 48 Cal. Rptr. 646, 651 (Ct. App. 1966). A challenge is “exercised when the challenged judge transfers the case for reassignment,”102Truck Ins. Exch. v. Superior Ct., 78 Cal. Rptr. 2d 721, 724 (Ct. App. 1998) (permitting a party to file a second peremptory challenge because the first peremptory challenge was against the first judge who denied the motion and thereafter retired, rendering the issue moot). and it cannot be rescinded, no matter what—the dismissal of the movant makes no difference.103See Louisiana-Pacific Corp. v. Philo Lumber Co., 210 Cal. Rptr. 368, 369 (Ct. App. 1985). If no other judge is available, the disqualified judge should contact the Chairman of the Judicial Council to solicit the assignment of an outside judge.104Nail v. Osterholm, 91 Cal. Rptr. 908, 911 (Ct. App. 1970). In order to prevent the appearance of judicial impropriety, if the judge is assigned to more than one case concerning the same movant, they are disqualified from all such cases.105Woods v. Superior Ct., 235 Cal. Rptr. 687, 687–88 (Ct. App. 1987). A writ of mandate petition is the “exclusive means of appellate review” for the motion, irrespective of its success.106In re Sheila B., 23 Cal. Rptr. 2d 482, 485 (Ct. App. 1993). This process serves “judicial economy and fundamental fairness” by “eliminat[ing] the waste of time and money which inheres if the litigation is permitted to continue unabated, only to be vacated on appeal because the subsequent rulings and judgments were declared ‘void’ by virtue of the erroneously denied disqualification motion.” Id. Upon a failed motion, the litigant has two avenues of redress before appeal: review by a different judge, like the district’s chief judge, and mandamus review.107Jeffrey W. Stempel, Judicial Peremptory Challenges as Access Enhancers, 86 Fordham L. Rev. 2263, 2269 (2018). On the other hand, litigants who never assert a challenge will have “forfeited the right to complain about [how the trial court’s alleged bias affected subsequent rulings] on appeal.”108People v. Lewis, 140 P.3d 775, 798 (Cal. 2006); see Mueller v. Chandler, 31 Cal. Rptr. 646, 647 (Ct. App. 1963).

2.  Judicial Rules

Section 170.6 intersects with other bodies of judicial rules, complicating the tapestry of California peremptory disqualification law. Upholding impartiality in the courts permeates all the guidelines that judges should follow, regardless of origin—Standard 10.20(b)(3) of the California Rules of Court instructs judges to “ensure that all orders, rulings, and decisions are based on the sound exercise of judicial discretion and the balancing of competing rights and interests and are not influenced by stereotypes or biases.”1092023 Cal. Rules of Ct. § 10.20(b)(3) (Jud. Couns. of Cal. 2023). In light of this goal, the California Rules of Court encourage outreach to the community110Id. § 10.20(a) (“[E]ach court should work within its community to improve dialogue and engagement with members of various cultures, backgrounds, and groups to learn, understand, and appreciate the unique qualities and needs of each group.”); Id. § 10.20(c)(3) (“Each committee should . . . [e]ngage in regular outreach to the local community to learn about issues of importance to court users.”). and collaboration with local committees and bar associations that endorse programs designed to educate about unconscious biases.111Id. § 10.20(c)(2). When litigants encounter judges who ignore Standard 10.20 of the California Rules of Court, they can submit complaints of bias either directly to the court or to the Commission on Judicial Performance without losing their statutory remedy through Section 170.6.112Id. § 10.20(d). California Code of Judicial Ethics Canon 3D(4), Government Code section 68725, and Rule 104 of the Rules of the Commission on Judicial Performance obligates judges to cooperate with the Commission on Judicial Performance.113Cal. Code of Jud. Ethics Canon 3D(3), cmt. 3D(4) (Cal. Judges Ass’n 2015). In one instance, the Commission on Judicial Performance ordered the removal of a judge who communicated with the potential movant to stop their challenge.114Inquiry Concerning Laettner, 8 Cal. 5th CJP Supp. 1, 54 (Comm. on Jud. Performance 2019).

California Code of Judicial Ethics Canon 3B(5), Canon 3C(1), and Canon 3E(5)(f)(iii), among others, comport with the Model Code of Judicial Conduct rule 2.3115Model Code of Jud. Conduct r. 2.3 (Am. Bar Ass’n 2020). because they advise that judges should be free of bias.116Cal. Code of Jud. Ethics Canon 3B(5), 3C(1), and 3E(5)(f)(iii) (Cal. Judges Ass’n 2015). Similar to the California Standards of Judicial Administration, California Code of Judicial Ethics Canon 3B(6) directs judges to “require lawyers in proceedings before [them] to refrain from . . . bias.”117Id. at Canon 3B(6). They consequently may feel cognitive dissonance (psychological discomfort from simultaneously complying with incongruous beliefs118Cognitive Dissonance, Merriam-Webster (Dec. 3, 2022), https://www.merriam-webster.com/dictionary/cognitive%20dissonance [https://perma.cc/VKG8-VVUW].) from essentially allowing litigants to discriminate against them using Section 170.6. Granted, these “standards, insofar as they may conflict with [S]ection 170.6, would be ‘invalid’ since the Judicial Council may only make rules which are not inconsistent with statute,”119People v. Superior Ct., 10 Cal. Rptr. 2d 873, 879 (Ct. App. 1992). but this paradox may still trouble them. Furthermore, California Code of Judicial Ethics Canon 3D(1) instructs judges with reliable information on the violations of other judges to report those violations to the appropriate authority and take any other corrective actions.120Jud. Ethics Canon 3D(1). The advisory committee’s commentary states that “[a]ppropriate corrective action could include direct communication with the judge or lawyer who has committed the violation, writing about the misconduct in a judicial decision, or other direct action, such as a confidential referral to a judicial or lawyer assistance program, or a report of the violation to the presiding judge, appropriate authority, or other agency or body.” Id. at Canon 3(D)(2). Considering these numerous regulations that either officially or informally punish judges who harbor biases, whether explicit or not, are peremptory disqualifications truly necessary?121Incentives are vital to eliminating bias in both judges and jurors. See Suzy J. Park, Racialized Self-Defense: Effects of Race Salience on Perceptions of Fear and Reasonableness, 55 Colum. J.L. & Soc. Probs. 541, 571 (2022) (“[S]ince the data suggest that it is difficult to make people ‘turn off’ their prejudices through the use of race salience, it is critical to choose jurors who are internally and genuinely motivated to be unprejudiced.”).

3.  Comparison to Peremptory Juror Challenges

In California, judicial peremptory challenges enjoy less resistance than in some other jurisdictions122See, e.g., Miller-El v. Dretke, 545 U.S. 231, 272 (2005) (Breyer, J., concurring) (criticizing the peremptory challenge system as a whole); Swain v. Alabama, 380 U.S. 202, 244 (1965) (Goldberg, J., dissenting) (“Were it necessary to make an absolute choice between the right of a defendant to have a jury chosen in conformity with the requirements of the Fourteenth Amendment and the right to challenge peremptorily, the Constitution compels a choice of the former.”); State v. Veal, 930 N.W.2d 319, 480 (Iowa 2019) (“[T]he only way to stop the misuse of peremptory challenges is to abolish them.”); Minetos v. City Univ. of N.Y., 925 F.Supp. 177, 183 (S.D.N.Y. 1996) (“[A]ll peremptory challenges should now be banned as an unnecessary waste of time and an obvious corruption of the judicial process.”). but are nonetheless more controversial than peremptory juror challenges. The California legislature passed AB 3070 in 2020—a proposal to require “the party exercising the peremptory challenge [to] show by clear and convincing evidence that an objectively reasonable person would view the rationale as unrelated to a prospective juror’s race, ethnicity, gender, gender identity, sexual orientation, national origin, or religious affiliation.”123Cal. State Assemb. 3070, 2020 Leg., 2019–2020 Reg. Sess. (Cal. 2020) (emphasis added); see Brian T. Gravdal, AB 3070 and Peremptory Juror Challenges in California: Strengthening Protection Against Discriminatory Exclusion, Berman Berman Berman Schneider & Lowary LLP, https://b3law.com/all-cases-list/ab-3070-and-peremptory-juror-challenges-in-california [https://perma.cc/6B6Q-ELDK]. The Legislature deliberately replaced the need to show purposeful discrimination under an objective standard in order to better target unconscious bias.124Gravdal, supra note 123. Unlike judicial peremptory disqualification, in which there is confusion regarding who holds the cause of action for discriminatory exclusion,125Infra p. 281. the bill clearly gives the right to both the party and the trial court. California Governor Gavin Newsom signed the legislation into law (California Code of Civil Procedure § 231.7), which went into effect for criminal trials on January 1, 2022 and will take effect for civil trials starting January 1, 2026.126Cal. Civ. Proc. § 231.7 (Deering 2023). The statute joined reforms in other states:127See Batson Reform: State by State, Berkeley L., https://www.law.berkeley.edu/experiential/clinics/death-penalty-clinic/projects-and-cases/whitewashing-the-jury-box-how-california-perpetuates-the-discriminatory-exclusion-of-black-and-latinx-jurors/batson-reform-state-by-state [https://perma.cc/3L46-M5ZQ]. Washington enacted a similar procedure in 2018 that was praised as a solution to Batson v. Kentucky,128See Daniel Edwards, The Evolving Debate Over Batson’s Procedures for Peremptory Challenges, Nat’l Ass’n of Att’ys Gen. (Apr. 14, 2020), https://www.naag.org/attorney-general-journal/the-evolving-debate-over-batsons-procedures-for-peremptory-challenges [https://perma.cc/EQ49-8CZ5]; Am. Soc’y of Trial Consultants, ASTC Position Paper on the Elimination of Peremptory Challenges: And Then There Were None 16 (2022), https://www.astcweb.org/resources/Documents/ASTC%20Position%20Paper%20on%20the%20Elimination%20of%20Peremptory%20Challenges%20-%20FINAL%207-14-2022.pdf [https://perma.cc/6VAH-T4W2]. a landmark case prohibiting unconstitutional discrimination during jury selection.129Batson v. Kentucky, 476 U.S. 79, 99 (1985) (“By requiring trial courts to be sensitive to the racially discriminatory use of peremptory challenges, our decision enforces the mandate of equal protection and furthers the ends of justice.” (footnote omitted)); see Jim Frederick, New Jury Selection Procedure in California: Is This the End of Peremptory Challenges? Is This the End of Batson?, Nat’l L. Rev. (Dec. 2, 2020), https://www.natlawreview.com/article/new-jury-selection-procedure-california-end-peremptory-challenges-end-batson [https://perma.cc/3YMF-DJDE] (“[Batson] requires a prima facie case of discrimination to be made before a party must explain the exclusion of a prospective juror by offering a facially neutral justification for the strike.”). After observing Batson’s shortcomings in actually resolving racial bias and discrimination,130See, e.g., Paula Hannaford-Agor, The Changing Civil Jury: Getting Back to “Normal”: Jury Trials in the Post-Covid Era, 98 Advocate 38, 40 (2022) (“[M]ost judges and lawyers privately agree that the Batson framework has been ineffective at curbing discrimination in jury selection.”); Gregg Costa, A Judge Comments, 48 Litigation 36, 36 (2022) (“According to a study aptly titled Thirty Years of Disappointment, as of 2016, North Carolina appellate courts had never found that a prosecutor violated Batson!”). See generally Anna Offit, Race-Conscious Jury Selection, 82 Ohio St. L.J. 201 (2021) (reporting a study of Assistant U.S. Attorneys showing how prosecutors consider race when striking jurors due to Batson). supporters argue that the spirit and letter of this legislative decision carries the promise of giving life to the federal precedent.131See La Rond Baker, Salvador A. Mungia, Jeffrey Robinson, Lila J. Silverstein, & Nancy Talner, Fixing Batson, 48 Litigation 32, 38 (2022); Robert Gavin, Chief Judge Highlights Proposal to Weed Out ‘Unconscious Racism’ on Juries, Times Union (Aug. 16, 2022), https://www.timesunion.com/news/article/Chief-judge-highlights-proposal-to-weed-out-17374759.php [https://perma.cc/Q6ET-PEXE]; Reforms Addressing Jury Selection Bias Proposed in New York and New Jersey, Equal Just. Initiative (Aug. 25, 2022), https://eji.org/news/reforms-addressing-jury-selection-bias-proposed-in-new-york-and-new-jersey [https://perma.cc/YQB9-URNE].

But this law is not without dissenters. The Alliance of California Judges rebukes it for creating “confusion and delay” since “lawyers could challenge every peremptory challenge made by the other side.”132Jim Frederick, New Jury Selection Procedure in California: Is This the End of Peremptory Challenges? Is This the End of Batson?, Faegre Drinker on Products (Dec. 2, 2020), https://www.faegredrinkeronproducts.com/2020/12/new-jury-selection-procedure-in-california-is-this-the-end-of-peremptory-challenges-is-this-the-end-of-batson [https://perma.cc/37UD-FBH8]. Coburn R. Beck, The Current State of the Peremptory Challenge, 39 Wm. & Mary L. Rev. 961, 1000 (1998) also stresses the restoration of peremptory juror challenges to its traditional form, not Batson-like modifications that can “produce[] a [confusing] circuit split over a trial procedure firmly established since the beginning of our nation.” Ultimately, it falls short of putting an end to peremptory juror challenges altogether, making others think that the change is not enough: Senate Bill 212 was introduced in 2021, which would have abolished peremptory challenges in criminal cases.133S. 212, 2021–2022 Leg., Reg. Sess. (Cal. 2021). One wonders at the end of the day if it is still accurate to classify this challenge as peremptory as “[a] challenge subject to questioning and explanation is, by definition, not peremptory.”134Beck, supra note 132, at 997–98. Whether judicial peremptory challenges can inherit this reform such that it is workable to judges poses an interesting question.

C.  Comparison Between California Law and Other States’ Law

Section 170.6 is one of the two forms of judicial peremptory challenges practiced by twenty states. Judicial officers in thirteen states, like jurors, can be substituted upon request without any accusation of improper personal interest, while those in the remaining seven states can only be substituted upon an affidavit of bias.135The states are Alaska, Arizona, California, Hawaii, Idaho, Illinois, Indiana, Kansas, Minnesota, Missouri, Montana, Nevada, New Mexico, North Dakota, Oregon, South Dakota, Texas, Washington, Wisconsin, and Wyoming. Gary L. Clingman, A Clash of Branches: The History of New Mexico’s Judicial Peremptory Excusal Statute and a Review of the Impact and Aftermath of Quality Automotive Center, LLC v. Arrieta, 46 N.M. L. Rev. 309, 336–37 (2016); see Alaska Stat. § 22.20.022 (LexisNexis 2023); Ariz. Rev. Stat. § 12-409 (LexisNexis 2023); Cal. Civ. Proc. Code § 170.6 (Deering 2023); Haw. Rev. Stat. Ann. § 601-7 (LexisNexis 2023); Idaho Code § 40(d)(1) (LexisNexis 2023); 725 Ill. Comp. Stat. Ann. 5/114-5(a) (LexisNexis 2023); Ind. Code § 35-36-5-1 (2023); Kan. Stat. Ann. § 20.311(d) (LexisNexis 2023); Minn. Stat. § 542.16 (2023); Mo. R. Civ. Pro. § 51.05 (LexisNexis 2023); Mont. Code Ann. § 3-1-804 (West 2023); Nev. Rev. Stat. Ann. § 1.230 (West 2023); N.M. Stat. Ann. § 38-3-9 (2023); N.D. Cent. Code § 29-15-21 (2023); Or. Rev. Stat. § 14.260 (West 2023); S.D. Codified Laws § 15-12-22 (LexisNexis 2023); Tex. Gov’t Code Ann. § 74.053 (LexisNexis 2023); Wash. Rev. Code Ann. § 4.12.040-50 (LexisNexis 2023); Wis. Stat. § 801.58 (LexisNexis 2023); Wyo. R. Civ. P. 40.1(b)(1). Table 2 below describes some differences between Section 170.6 and peremptory challenge statutes in other states:

Table 2.
ParametersSection 170.6Statutes in Other States
Is there a fee for reassignment to a new judge?NoYesa
How many challenges to a party per case?OneTwob
Does alignment of interest or lack thereof define a party (or side)?YesNoc
Must litigants informally ask the judge to voluntarily recuse from the case before filing an affidavit?NoYesd
Can criminal litigants transfer their case to a different judge?YesNoe
Sources:  a  See, e.g., Mont. Code Ann. § 3-1-804 (West 2023); Nev. Rev. Stat. Ann. § 1.230 (West 2023). b  E.g., Or. Rev. Stat. § 14-250-70 (West 2023). c  Mo. R. Civ. Pro. § 51.05 “divides the parties into classes (e.g. plaintiffs, defendants, third party plaintiffs, third party defendants, interveners) and affords one change of judge per class” and Nev. Rev. Stat. Ann. § 1.230 treats “[e]ach action, whether single or consolidated . . . as having only two sides. Clingman, supra note 135, at 337–38. d  S.D. Codified Laws § 15-12-22 (2023); see Clingman, supra note 135, at 338. e  Ind. Code Ann. § 35-36-5-1, Nev. Rev. Stat. Ann. § 1.230, Tex. Gov’t Code Ann. § 74.053, and Wyo. R. Civ. P. 40.1(b)(1) are some statutes that recognize this right in civil cases only. Clingman, supra note 135, at 338.

Although these states are not uniform in protocol, they all place weight on a movant’s good faith and decline to investigate whether the movant’s reasons, if even stated, are true, differentiating state law from federal law, which demands supporting facts.

II.  THE POLICY TRADE-OFFS

A.  Public Confidence in the Judiciary

Those who applaud the judicial peremptory challenge, a device that makes it easier to disqualify judges, emphasize its utility in maintaining and increasing public confidence that the judiciary will deliver equal justice under the law. Although “the law, not any individual or group, is a judge’s only legitimate constituent,” judges have free speech protections in judicial election campaigns.136Thomas R. Phillips & Karlene Dunn Poll, Free Speech for Judges and Fair Appeals for Litigants: Judicial Recusal in a Post-White World, 55 Drake L. Rev. 691, 694 (2007); David K. Stott, Zero-Sum Judicial Elections: Balancing Free Speech and Impartiality Through Recusal Reform, 2009 BYU L. Rev. 481, 481 (2009) (arguing that judicial candidates have a First Amendment right to express their opinions to the electorate and receive campaign contributions—hence, judicial elections create a zero-sum game). If their views on controversial legal and political issues are broadcast through various media outlets, the public will naturally lose hope that due process137Marshall v. Jerrico, Inc., 446 U.S. 238, 242 (1980) (“The Due Process Clause entitles a person to an impartial and disinterested tribunal in both civil and criminal cases.”); see Procedural Due Process Civil, Justia, https://law.justia.com/constitution/us/amendment-14/05-procedural-due-process-civil.html [https://perma.cc/C9KT-QPBM]. Judge Friendly argued that an unbiased tribunal is indispensable to due process. Peter Strauss, Due Process, Legal Info. Inst. (Oct. 2022), https://www.law.cornell.edu/wex/due_process [https://perma.cc/2AWK-UMJL]. still exists in the courtroom. Since judges will likely reveal their biases, there are practitioners who advocate for appellate courts to adopt the peremptory strike system, as in California trial courts through Section 170.6, so the public can trust that their matters will be heard by a neutral arbitrator.138Phillips & Poll, supra note 136, at 718–20. They echo Justice Kennedy’s advice for states to “adopt[] recusal standards more rigorous than due process requires” in an effort to protect judicial integrity.139Republican Party of Minn. v. White, 536 U.S. 765, 794 (2002); see Serbulea, supra note 45, at 1146 (“Recusal motions are different than other procedural motions because they implicate the very legitimacy of the legal system.”). Meanwhile, worried about the increasing caseload burdening the federal judiciary, some academics urge Congress to set up a commission responsible for establishing a judiciary reform act that would go into effect in 2030.140Peter S. Menell & Ryan Vacca, Revisiting and Confronting the Federal Judiciary Capacity “Crisis”: Charting a Path for Federal Judiciary Reform, 108 Cal. L. Rev. 789, 879 (2020). As the judiciary becomes more congested with inefficient case management and reduced dockets, judicial competence suffers, especially considering the “10% problem,” which is a “rough estimate of the percentage of district court judges who are considered unfit or limited in their capacity to dispense justice fairly.”141Id. at 884–85. These academics provide the 2030 Commission with a solution: peremptory challenges.142Id. at 885.

Curiously, the reason cited for condemning the challenge sounds familiar—increasing public confidence in the administration of justice. One legal scholar believes peremptory disqualification injures the judiciary’s reputation because “automatic transfer does not permit a judge to refute the allegations of bias, and so may create the public impression that more judges are biased, or have conflicts of interests, than is actually the case.”143Frost, supra note 40, at 587; see Serbulea, supra note 45, at 1144 (“Allowing peremptory challenges will most likely result in an increased number of disqualifications.”). Regrettably, judicial discretion, in which “the law gives the judge a range of options and choices, or relies on the judge’s assessment of the circumstances in drawing further conclusions,” exposes judges to criticism without crisis managers to guide them through this era of social media.144Levi, supra note 23. As committees and organizations dedicated to judicial independence face extinction, many stress the need for the legal profession to rally in defense of judges, perhaps by devoting resources to educating the public on what judges actually do.145See, e.g., id. (“It is distressing that in recent years we have seen the demise of two leading organizations most devoted to judicial independence—the American Judicature Society and Justice at Stake—as well as the defunding of the one American Bar Association committee dedicated to judicial independence.”); Serbulea, supra note 45, at 1149 (“Educating the people about the judicial system and its inner workings will increase the public’s confidence in the judicial system.”).

B.  Abuse of the Challenge

Peremptory challenges to judges can disrupt the harmony between not just a litigant and their judge, but also the litigant’s attorney and the judge, as well as the litigant and their attorney, essentially poisoning the most material relationships in the courtroom. First, the litigant must present the motion to the very judge they want disqualified.146E.g., Lewis v. Linn, 26 Cal. Rptr. 6, 9 (Ct. App. 1962). Since the judge knows the movant’s identity, they may feel “frustrated at being required to grant relief to a party who had made what [they] consider to be an unwarranted slight to their integrity.”147Geoffrey P. Miller, Bad Judges, 83 Tex. L. Rev. 431, 481 (2004). If the motion is rejected, the litigant is stuck with the allegedly biased (and now insulted) judge until appeal because it is difficult to prevail on other review proceedings.148See Stempel, supra note 107, at 2269. Second, although the attorney might wish to evade a particular judge for their entire legal career, a successful motion in one case does not insulate them from the judge’s hostility in future cases.149Miller, supra note 147, at 481–82 (“While litigants may never appear in the judge’s courtroom again, the attorney probably will, and judges have long memories. Judges may bide their time and then take out their frustration on an attorney in another case.”). Third, caught in a web of ethical obligations, the attorney deals with an uncomfortable dilemma: Are they loyal to the judge or their client? No matter their self-interest to stay on good terms with the judge, they must reconcile their duty of vigorous advocacy on behalf of their client with their duty of honesty and respect to the court. They are probably tempted to use their affidavit power (that is, to capitalize on the boilerplate affidavit requesting only a conclusory accusation of bias) to win their client’s case, for they are given the benefit of the doubt.150For explanatory hypotheticals, see Miller, supra note 147, at 482. This temptation is why “allow[ing] peremptory challenges only on consent of both parties with the challenges waived if no agreement is reached,” a proposal to remedy peremptory juror challenges, would not work, at least in the context of judicial disqualification. Caren Myers Morrison, Negotiating Peremptory Challenges, 104 J. Crim. L. & Criminology 1, 7 (2014). However, their capability as “true advocates” is impeded by ethical rules that impose professional discipline should they lie about judges.151See, e.g., Model Rules of Pro. Conduct r. 8.2(a) (Am. Bar Ass’n 2023) (“A lawyer shall not make a statement that the lawyer knows to be false or with reckless disregard as to its truth or falsity concerning the qualifications or integrity of a judge . . . .”); Model Rules of Pro. Conduct r. 8.4(d) (Am. Bar Ass’n 2023) (“It is professional misconduct for a lawyer to . . . engage in conduct that is prejudicial to the administration of justice . . . .”). They also cannot claim the full extent of free speech rights under the First Amendment, presumably fueling their apprehension at the growing number of sanctions in 2022.152See generally John B. Harris, Lawyers Beware: Criticizing Judges Can Be Hazardous to Your Professional Health, Frankfurt Kurnit Klein + Selz PC (Feb. 1, 2022), https://professionalresponsibility.fkks.com/post/102hhmt/lawyers-beware-criticizing-judges-can-be-hazardous-to-your-professional-health [https://perma.cc/YCZ6-YGPZ] (discussing both new and old cases regarding attorneys’ criticism of judges to demonstrate a trend toward discipline).

Since the genesis of peremptory disqualification statutes, the risk of “judge shopping” has haunted legal scholars and practitioners alike. They argue that marginal improvements to judicial accountability do not warrant sacrificing judicial independence and integrity. When litigants judge shop under the guise of eliminating bias, they perpetuate the narrative that judges are simply “politicians in black robes” even though the Model Code of Judicial Conduct, court rules, judicial discipline sanctions,153See generally Cynthia Gray, A Study of State Judicial Discipline Sanctions (2002). and public opinion motivate judges to act properly. Admittedly, this illegitimate purpose is not allowed; the challenge, however, is an absolute right without regard for pretenses. For instance, according to one columnist, “[Section 170.6] could be warranted against the judge who tends to let all of [their] cases go to trial” if the attorney is “hoping to escape [the case] via summary judgment.”154Rick Merrill, Tech Tip: Using Judicial Analytics to Stay One Step Ahead, 60 Orange Cnty. Law. 50, 50–51 (2018).

To rebut these complaints of abuse, the challenge’s supporters point to the stringent rules governing the motion: specifically, its timing (framed as rushing litigants to “move[] as expeditiously as . . . is possible . . . after theretofore agreed on matter becomes litigated”155Mayr v. Superior Ct., 39 Cal. Rptr. 240, 242 (Ct. App. 1964).) and form. These restrictions should discourage litigants from not only judge shopping, but also “from waiting to see how the judge views the case and rules on motions before making the peremptory challenge decision.”156Stempel, supra note 107, at 2273–74. In People v. Rojas, 31 Cal. Rptr. 417, 420 (Ct. App. 1963), the defendants peremptorily challenged the judge over three years after judicial assignment when the judge had already heard their case and found them guilty. Section 170.6, for example, is a “limited right and is not a vehicle for disqualifying judges in all situations in which there is a potential for bias.”157Matthews v. Superior Ct., 42 Cal. Rptr. 2d 521, 524 (Ct. App. 1995) (emphasis added). One law professor advances a limited conception of misuse such that litigants (1) can avoid extremist judges who are not necessarily biased but (2) cannot technically judge shop since a randomly assigned judge will preside over the previously assigned judge’s disqualification.158Stempel, supra note 107, at 2274–75 (“For example, a defense attorney may want to eject a harsh sentencing ‘hanging’ judge from the case . . . . But it hardly makes the challenge improper when used to avoid judges at the extremes in terms of both jurisprudential tendencies and competence.”). Peremptory disqualification increases the chance of the new judge sharing the same beliefs as most judges, which promotes a representative judiciary that reflects the citizenry because “an average judge may be more representative than a random one.”159John Leubsdorf, Theories of Judging and Judge Disqualification, 62 N.Y.U. L. Rev. 237, 273 (1987). This kind of judge shopping, as defined by the challenge’s opponents, is akin to “forum shopping,”160Stempel, supra note 107, at 2275–76 (“For example, litigants may employ the following strategies: removal to federal court; a “minimum contacts” approach to personal jurisdiction; a liberal approach to venue (but subject to the possibility of transfer to a more convenient venue); stringent enforcement of forum selection and choice of law clauses, including arbitration or other forum-specific dispute-resolution clauses; and clever selection of particular plaintiffs or claims in order to bring a test case or a potentially precedent-setting case in a favorable forum.” (footnotes omitted)). but the former is attacked as a radical threat to American ideals, while the latter enjoys more forgiveness from critics. The same can be said of “filing several cases simultaneously and dismissing all but the case before one’s preferred judge.”161Nancy J. King, Symposium on Race and Criminal Law: Batson for the Bench? Regulating the Peremptory Challenge of Judges, 73 Chi.-Kent L. Rev. 509, 523 (1998). Besides statutory safeguards, like the very short window of opportunity to exercise a challenge, litigants might eschew the challenges—if their motion succeeds, they risk an even more unfavorable judicial draw, but if their motion fails, they risk a resentful judge.

C.  Intimidation of Judges

There is also a strong assertion that judges will encounter intimidation, further cementing the deadlock between the two stances. A judge is more likely to be influenced by pressure from the litigant and their attorney when the defendant’s life, liberty, and property are hanging in the balance—in other words, criminal cases. Does the judicial peremptory challenge enable prosecutors to shop for “law and order” judges who are “tough on crime”?162Per a study of San Diego courts in the late 1970s, district attorneys used the challenge against defendant-friendly judges. Pamela J. Utz, Settling the Facts: Discretion and Negotiation in Criminal Court 78, 84 (1978). If a judge is peremptorily disqualified from every criminal matter to which they are assigned (colloquially known as “papering” or “blanket challenges”),163See, e.g., Roger M. Grace, Gascón Crosses the Line—Again, Metro. News–Enter. (May 3, 2022), http://www.metnews.com/articles/2022/PERSPECTIVES_050322.htm [https://perma.cc/GMF9-882M] (reporting that a head deputy District Attorney instructed all deputy District Attorneys to file a disqualification motion under section 170.6 every time a case was assigned to a certain judge in 2022); Dakota Morlan, Calaveras County DA ‘Papering’ Superior Court Judge with Disqualifications, Calaveras Enter. (May 7, 2021), https://www.calaverasenterprise.com/articles/crime/calaveras-county-da-papering-superior-court-judge-with-disqualifications [https://perma.cc/4C9B-ED8P] (reporting that the Calaveras County District Attorney’s office filed dozens of peremptory challenges against a single judge within ten days in 2021). they risk not only a non-criminal reassignment that poses a “very serious problem for a judge whose entire legal career has been spent in the criminal justice system,”164James Michael Scheppele, Are We Turning Judges into Politicians?, 38 Loy. L.A. L. Rev. 1517, 1524 (2005). but also transfer to a court that is, in their opinion, more inconvenient or less prestigious.165Ted Rohrlich, Scandal Shows Why Innocent People Plead Guilty, L.A. Times, Dec. 31, 1999, at A1 (“If you called the police liars, they’d [issue a peremptory challenge against] you . . . . [I]nstead of working on a nice assignment near your home, they [your fellow judges] send you downtown or to juvenile or dependency court, where they send the slugs.”). As judges try to appease prosecutors to avoid repeated disqualification, the pool of judges actually deciding criminal cases becomes undersaturated with lenient and liberal judges.166Adam Peterson, The Future of Bail in California: Analyzing SB 10 Through the Prism of Past Reforms, 53 Loy. L.A. L. Rev. 263, 268 (2019). Prosecutorial control over judges (along with other challenges due to unusual judicial philosophies) results in a much smaller spectrum of worldviews among judges, hampering the development of legal interpretations and encumbering healthy debate.167For an article discussing how peremptory juror challenges make it more probable that the jury will be composed entirely of jurors on one extreme of an ideological spectrum, see Francis X. Flanagan, Peremptory Challenges and Jury Selection, 58 J.L. & Econ. 385, 385 (2015). One law professor argues that the challenges hurt jurors’ ability to render accurate verdicts by “systematically eliminating jurors with a range of perspectives who might have challenged erroneous or mistaken ideas.” Nancy S. Marder, Beyond Gender: Peremptory Challenges and the Roles of the Jury, 73 Tex. L. Rev. 1041, 1045 (1995). Erwin Chemerinsky responds to the “unlimited use of peremptory challenges against a single judge, albeit in different cases,” by prescribing even greater procedural protections.168Laurie L. Levenson, The Rampart Scandal: Policing the Criminal Justice System: Unnerving the Judges: Judicial Responsibility for the Rampart Scandal, 34 Loy. L.A. L. Rev. 787, 812–13 (2001) (commenting on Erwin Chemerinsky, An Independent Analysis of the Los Angeles Police Department’s Board of Inquiry Report on the Rampart Scandal, 40 Loy. L.A. L. REV. 545 (2001)).  However, even if litigants manipulate this mechanism to pressure judges,169Scheppele, supra note 164, at 1523. it is hard to imagine judges succumbing to partiality after only one or even a few cases. Perhaps district attorneys or public defenders who frequently appear in the same court can effectively intimidate judges,170See id. at 1523–24. but in the big picture, criminal cases make up a small subset of total filings.171In 2022, for instance, there were 309,102 civil filings and 71,111 criminal filings in the U.S. district courts. Federal Judicial Caseload Statistics 2022, U.S. Cts., https://www.uscourts.gov/statistics-reports/federal-judicial-caseload-statistics-2022 [https://perma.cc/W74D-NVX6].

D.  Discrimination Against Judges

Two law professors illustrate how judicial peremptory challenges can act as a vehicle for discrimination: one uses a hypothetical,172Jack H. Friedenthal, Exploring Some Unexplored Practical Issues, 47 St. Louis L.J. 3, 9 (2003) (“Suppose that an employment discrimination case is filed by a woman in a state court which has an automatic dismissal law, against a handful of male defendants with related yet somewhat factually divergent interests that, at least technically, are hostile to one another. The pool of judges available to try the case consists of a number of females whom the lawyers for the defendants fear may tend to favor plaintiff’s case. Suppose further that counsel for each of the defendants agrees that each, in turn, will automatically eliminate any female judge who is initially or subsequently assigned to try the case, thus virtually ensuring that a male judge will ultimately be selected.”). while the other uses two cases in which attorneys were accused of discriminating against their judges.173King, supra note 161, at 512–13 (summarizing People v. Williams, 54 Cal. Rptr. 2d 521 (Ct. App. 1996), in which the prosecution’s peremptory challenge against a Black judge in a case concerning two Black criminal defendants was scorned by the public as racist, and People v. Williams, 774 P.2d 146 (Cal. 1989), in which a Black judge rejected a race-based peremptory challenge against him). There is a 1985 study suggesting that race-based abuse of the challenge was rare,174See Larry C. Berkson & Sally Dorfmann, Judicial Substitution: An Examination of Judicial Peremptory Challenges in the States 142 tbl.VII-9 (1986) (reporting the 1985 study’s findings—among those surveyed, 10% of defense attorneys, 4% of chief judges, and 1% of prosecutors thought judges were peremptorily disqualified due to race). but the latter professor dismisses its applicability because there is now greater awareness about the unconstitutionality of racially charged decisions,175Consider the Black Lives Matter and Anti-Asian Hate movements that shed light on racial inequality. See Hannaford-Agor, supra note 130, at 39 (“Within weeks of George Floyd’s murder, dozens of state-court systems had convened task forces and commissions charged with identifying the root causes and drafting recommendations to address the lack of demographic diversity in jury pools and juries.”). especially after Batson and J.E.B. v. Alabama ex rel. T.B.176J.E.B. v. Ala. ex rel. T.B., 511 U.S. 127, 146 (1994) (“When persons are excluded from participation in our democratic processes solely because of race or gender, this promise of equality dims, and the integrity of our judicial system is jeopardized.”). With more judges who identify with marginalized groups, there are consequently more opportunities for challenges based on protected characteristics (leading to disproportionate disqualifications along racial lines, for example).177King, supra note 161, at 517 (“Because the bench has consisted almost entirely of white judges until the last several years, only recently have litigants had the ability to shop for a judge of a particular race or ethnicity. In particular, there were very few, if any, judges of color on the bench in the predominantly western and mid-western states that authorized judicial peremptory challenges at the time when past studies were conducted.” (footnotes omitted)); see Mentoring Program Aims to Increase Diversity of Judge Applicants, Cal. Cts. Newsroom (Mar. 5, 2021), https://newsroom.courts.ca.gov/news/mentoring-program-aims-increase-diversity-judge-applicants (“For the 15th straight year, California’s judicial bench has grown more diverse . . . . [A] new mentorship program in Los Angeles County seeks to accelerate the diversity of the bench . . . .”). In California, where 63.1% of judges are white, a white judge will probably substitute a disqualified judge of color.178Jud. Couns. of Cal., supra note 27, at 1. These removals are contrary to a socioeconomically representative judiciary—judicial officers from historically oppressed groups are more likely to have public-interest experience and less likely to have a upper-class background than their colleagues.179King, supra note 161, at 521. Additionally, empirical studies implying that age and gender are outcome determinative may tempt litigants into issuing ageist or sexist challenges.180See, e.g., Morris B. Hoffman, Francis X. Shen, Vijeth Iyengar & Frank Krueger, The Intersectionality of Age and Gender on the Bench: Are Younger Female Judges Harsher with Serious Crimes?, 40 Colum. J. Gender & L. 128, 164 (2020) (“Younger female judges sentence high-harm cases significantly more harshly than their male and older female colleagues.”); Maureen A. Howard, Taking the High Road: Why Prosecutors Should Voluntarily Waive Peremptory Challenges, 23 Geo. J. Legal Ethics 369, 401 (2010) (“Ironically, research suggests that the two demographics that actually have some empirical validity (and are thus ‘rational’ bases for peremptories), are those that are specifically prohibited by the Constitution: race and gender.”). Other studies have confirmed the discriminatory effects of peremptory juror challenges,181See, e.g., C.J. Williams, Striking Some Strikes: A Proposal for Reducing the Number of Peremptory Strikes, 68 Drake L. Rev. 789, 817–18 (2020) (“The broader conclusion that can be reached from these studies is that the greater the number of peremptory strikes available to the parties, the less diverse the petit jury becomes regardless of the diversity of the jury venire.”). substantiating arguments that the challenge is inherently flawed and does more discriminatory harm than any good.182See, e.g., Alen v. State, 596 So.2d 1083, 1086 (Fla. Dist. Ct. App. 1992) (Hubbart, J., concurring) (“Rather than engage in a prolonged case-by-case strangulation of the peremptory challenge over a period of many years which in the end will effectively eviscerate the peremptory challenge or, at best, result in a convoluted and unpredictable system of jury selection enormously difficult to administer—I think the time has come, as Mr. Justice Marshall has urged, to abolish the peremptory challenge as inherently discriminatory.”); Morris B. Hoffman, Peremptory Challenges Should Be Abolished: A Trial Judge’s Perspective, 64 U. Chi. L. Rev. 809, 871 (1997) (“[E]ven assuming the peremptory challenge ever worked in this country as anything other than a tool for racial purity, and even assuming it is working today in its post-Batson configuration to eliminate hidden juror biases without being either unconstitutionally discriminating or unconstitutionally irrational, I submit that its institutional costs outweigh any of its most highly-touted benefits. Those costs—in juror distrust, cynicism, and prejudice—simply obliterate any benefits achieved by permitting trial attorneys to test their homegrown theories of human behavior on the most precious commodity we have—impartial citizens.”). This could ring true for judicial peremptory challenges as well: What is stopping attorneys who discriminate against jurors from also discriminating against judges?

Then again, litigants are forbidden from exercising the challenges solely based on group affiliation like race and ethnicity, gender, sexual orientation, religion, and so forth.183Peter David Blanck, The Appearance of Justice: The Appearance of Justice Revisited, 86 J. Crim. L. & Criminology 887, 903 (1996); see People v. Superior Ct., 10 Cal. Rptr. 2d 873, 884 (Ct. App. 1992) (“Section 170.6 cannot be employed to disqualify a judge on account of the judge’s race.”). They may not even want to rely on such factors—a judge’s “prior decisions made while on the bench, statements made in public forums, [and] professional and political reputations years deep”184King, supra note 161, at 521. But see Howard, supra note 180, at 401 (“Ironically, research suggests that the two demographics that actually have some empirical validity (and are thus ‘rational’ bases for peremptories), are those that are specifically prohibited by the Constitution: race and gender.”). better predict judicial propensity, after all. According to a member of the Alaska Judicial Council, the challenges in that state did not, in fact, depend on race or gender.185King, supra note 161, at 521 n.75. The problem, however, is not simply solved. Take California Code of Civil Procedure section 170.2, Section 170.6’s sister judicial disqualification statute, for example. It prohibits discrimination against judges, yet it does not seem to apply to Section 170.6.186Cal. Civ. Proc. Code § 170.2 (Deering 2023). Despite precedent that judges deserve shelter under the Equal Protection Clause of the Fourteenth Amendment,187See City of Cleburne v. Cleburne Living Ctr., 473 U.S. 432, 439 (1985) (“The Equal Protection Clause of the Fourteenth Amendment commands that no State shall ‘deny to any person within its jurisdiction the equal protection of the laws,’ which is essentially a direction that all persons similarly situated should be treated alike.” (citing Plyler v. Doe, 457 U.S. 202, 216 (1982)). “any party charging that [their] adversary has used a [S]ection 170.6 challenge in a manner violating equal protection bears the burden of proving purposeful discrimination,”188People v. Superior Ct., 10 Cal. Rptr. 2d 873, 884 (Ct. App. 1992). which is a high, if not unattainable, standard. Historical patterns of a movant’s discrimination could replace direct evidence of discriminatory intent, but there is a caveat: Is it possible to discern a pattern from a relatively small sample size?189King, supra note 161, at 524. Moreover, in the event the judge, as the right holder, declines to pursue their cause of action for discriminatory challenges, there is much uncertainty about whether litigants then have standing to object.190Id. at 528–32.

III.  EMPIRICAL FINDINGS IN CALIFORNIA

A.  Research Methodology

It appears that much of the policy debate about judicial peremptory disqualification is informed by theory rather than empirical data. Where are the surveys asking the public in states that allow the challenge about their confidence in the judiciary and perception of judicial bias? There is some research investigating how the challenges can intimidate judges (especially if initiated by prosecutors191See Utz, supra note 162, at 84 regarding the study of San Diego courts in the late 1970s and infra note 163 regarding District Attorneys’ offices and “papering” or “blanket challenges.”) and discriminate against judges of a certain race or gender.192See infra note 161 regarding the Alaska Judicial Council’s research and infra note 174 regarding the 1985 survey. Nonetheless, there remains a dearth of statistical findings regarding the frequency and type of abuses resulting from peremptory challenges in actual operation.193See, e.g., Miller, supra note 147, at 482 (“The frequency of peremptory challenges . . . do not appear to be maintained or distributed.”). “[W]ithout [the collection of empirical data], predictions about what attorneys will do [or cause] with peremptory challenges are guesswork,” leaving the aforementioned hypotheses with no answers.194King, supra note 161, at 515 n.49. Therefore, this Note aims to paint a more complete picture by tackling two questions: (1) do strict procedural rules really act as a barrier to slow the number of disqualifications, so the number of disqualified judges is roughly equal to the number of judges disciplined for bias, and (2) are there discriminatory effects based on judges’ political parties that prevent a representative judiciary? It will do so by adhering to the recommendations to examine orders on peremptory challenges in cases.195E.g., N.Y. State Just. Task Force, Recommendations Regarding Reforms to Jury Selection in New York 18–19 (2022) (“[E]xamination would likely take place through the creation of records . . . on peremptory challenges across cases, including tracking the stated reasons, if any, given for a challenge, and the judge’s ruling on the challenge.”).

Due to limited time and resources, only the available orders on LexisNexis (specifically, the 240 citing decisions of Section 170.6 after filtering for a timeline of January 1, 2021 to December 31, 2021) are analyzed. LexisNexis is a trustworthy source,196LexisNexis boasts the largest collection of caselaw, Products, LexisNexis, https://www.lexisnexis.com/en-us/products/lexis.page [https://perma.cc/3SHH-Z65F], with 1.2 million new legal documents added daily, About LexisNexis, LexisNexis, https://www.lexisnexis.com/en-us/about-us/
about-us.page [https://perma.cc/BRE6-WA5W], using a 29-step editorial process, Lexis Case Law Research by State, LexisNexis, https://www.lexisnexis.com/en-us/products/lexis/case-law-research.page [https://perma.cc/N9T8-BDCL].
but checking other databases, such as Westlaw, would have ensured that this methodology did not overlook orders. This truncated sample is largely not generalizable to the years before or after 2021. California is also not a microcosm for the entire nation: when interpreting the number of disqualified judges who were registered Democrats versus the number of disqualified judges who were registered Republicans, one should remember that California is a “blue” state that is considered a Democratic stronghold. Further, this Note interprets suggestive trends, not causal relationships, from the data as no formal statistical methods are used. Lastly, given that the cases’ dockets, including other related orders, opinions, and filings are not reviewed, there is missing information for some orders (for instance, the disqualified judge’s identity, the order’s date, whether the order was accepted or denied, and the reason behind the decision), which could misrepresent the results. For decisions that anonymously mention both the disqualified judge and the supervising or presiding judge, and hence create confusion about the role of the decision’s author, the analysis below errs on the side of caution and excludes these orders when tracking disqualified judges. Ideally, the study would only include orders from 2021; for consistency, it includes orders both without a date and from before 2021 if the decision that discusses the order is from 2021.

B.  Preliminary Empirical Data

There were 134 cases from January 8, 2021 to December 30, 2021 that revealed 158 ascertainable orders either accepting, denying, or discussing previously accepted or denied judicial peremptory challenges. For reference, there were 4,464,380 total filings in California superior courts in 2021.197Jud. Council of Cal., 2022 Court Statistics Report: Statewide Caseload Trends 78 (2022), https://www.courts.ca.gov/documents/2022-Court-Statistics-Report.pdf [https://perma.cc/M9FH-V6HZ]. Curiously, out of the 58 counties in California, only 11 (19%) had reported orders: Los Angeles (76 filed motions), Orange (41 filed motions), Sacramento (24 filed motions), San Diego (5 filed motions), Alameda (3 filed motions), Riverside (3 filed motions), Contra Costa (2 filed motions), Butte (1 filed motion), Madera (1 filed motion), San Francisco (1 filed motion), and Santa Clara (1 filed motion). Given that the study found 158 motions from just 134 cases and movants in only 11 out of 58 counties, this low occurrence of the challenges suggests that (1) litigants are generally not taking advantage of this litigation tool for improper purposes and (2) a majority of judges are perceived to be impartial. Since succeeding judges are randomly selected, there is a chance that litigants who detected bias in their judge would have issued a challenge if not for the fear that they might have to litigate under an even more biased judge. But, excluding pessimistic litigants who have little faith in judicial officers as a whole, it is unlikely that they will choose not to file a motion and endure a laborious litigation under a biased judge.

Among the 158 motions under Section 170.6, only 54 (34%) denied the challenge—Figure 1 displays the number of denials per specific reason:

Figure 1.

Section 170.6 is replete with rigorous procedural rules to make it harder for litigants to recklessly eliminate a qualified judge assigned to their matter. First, a little more than half of the denied motions (52%) failed to comply with Section 170.6(a)(2)’s timing standards.198E.g., Minute Order, Shurr v. Zuniga, No. 37-2018-00046744-CL-PA-CTL, 2021 Cal. Super. LEXIS 42241 (Dec. 10, 2021). Second, 4 challenges (7%) were defeated because they were addressed to an appellate judge and thus did not survive Section 170.6(a)(1).199E.g., Minute Order, Healy v. Orange Cnty. Super. Ct., No. 30-2021-01223007-CL-MC-CJC, 2021 Cal. Super. LEXIS 118829 (Sept. 29, 2021). Third, there was a three-way tie for reasons that blocked 3 motions (5.5%) each: submitting more than one challenge, in violation of Section 170.6(a)(4), whether it was from one party or one determined side;200E.g., Minute Order, Amezcua-Moll & Assoc. v. Modarres, No. 30-2017-00927161-CU-FR-NJC, 2021 Cal. Super. LEXIS 136812 (July 26, 2021). the continuation rule, codified by Section 170.6(a)(5); and the form standards, mandated by Section 170.6(a)(5) for written affidavits and Section 170.6(a)(6) for oral statements.201The three decisions for the continuation rule are Denny v. Arntz, No. A160234, 2021 Cal. App. Unpub. LEXIS 3104 (Cal. Ct. App., May 12, 2021); Minute Order, Cabral v. Walgreens Co., No. RG21093196, 2021 Cal. Super. LEXIS 64016 (July 13, 2021); and Order, Boesen v. Erickson, No. 20STCV36810, 2021 Cal. Super. LEXIS 98400 (Apr. 5, 2021). For the three motions that did not have proper form, some clarification may be helpful. One of the three orders is Minute Order, Simon v. Mercedes Benz United States, No. 30-2020-01157389, 2021 Cal. Super. LEXIS 106866 (Aug. 5, 2021), concerning a motion that did not address the correct judge. Another order is Order, Ramsey v. Uber Techs., No. MCC2000229, 2021 Cal. Super. LEXIS 142274 (Aug. 4, 2021), which ruled that the motion was not made under oath. The remaining order is Minute Order, Velasquez v. Doe #1, No. 30-2016-00833070-CU-PA-NJC, 2021 Cal. Super. LEXIS 27046 (Mar. 11, 2021), regarding a motion that did not name a judge at all. Fourth, 2 others (4%) were unsuccessful because the court had no authority.202E.g., People v. Moon, No. B306195, 2021 Cal. App. Unpub. LEXIS 5485 (Cal. Ct. App., Aug. 25, 2021). Fifth, 1 (2%) was denied as the movant had not yet appeared in the action.203Court Order, Second Site LLC v. Scott, No. BC723513, 2021 Cal. Super. LEXIS 73877 (Apr. 22, 2021). There was also a strange motion that was declared a “sham” since the litigant who filed the challenge was not a real person—as a result, the judge denied the challenge as it was not “duly presented” in accordance with Section 170.6(a)(4).204Minute Order, Hannaford v. Seven Satellite Pty, No. 19STCV13245, 2021 Cal. Super. LEXIS 76672, at *10 (July 30, 2021). Regrettably, the reasons for 9 denials (17%) could not be gleaned from the publicly available case material.205E.g., Order Resetting the Order to Show Cause Hearing for Why a Preliminary Injunction Should Not Issue and to Extend the Temporary Restraining Order, Genuis Fund I ABC v. Co. V, No. 20STCV39545, 2021 Cal. Super. LEXIS 25876 (May 28, 2021).

In line with the empirical finding that 66% of challenges were granted, attorneys, at least competent ones, are not only aware of the timing and form rules, but also successfully follow them. This comes as no surprise considering they are used to meeting the many deadlines that make up litigation. The odds of submitting a faulty challenge and consequently suffering under an offended judge are negligible—presumably, counsel would not carelessly file a motion that they know or should know is bound to fail. There are several safeguards that litigants must navigate, but untimeliness, the most frequent reason for rejected motions, was a weak barrier, stopping just 18% of the challenges. The statute counts on the time limit for filing the motion to prevent litigants from peremptorily disqualifying their judge based on how the judge has been ruling on the case. However, litigants or their attorneys may already know how the judge will view their case as soon as they receive the judicial assignment (or at least within the designated time frame) due to “random internet searches, anecdotal opinions from colleagues, or perhaps printed biographical material about the judge.”206Merrill, supra note 154, at 50. Hence, it looks like the limit on the number of challenges per case is the only statutory design that might effectively stall abuse, as litigants are reluctant to gamble that their new judge will not be worse.

The remaining 104 successful challenges (66%) disqualified at least 37 judges from 1 or more cases. Figure 2 below illustrates this proportion:

Figure 2.

Given there were 1,755 superior court judges in 2021,207State of Cal. Comm’n on Jud. Performance, 2021 Annual Report 11 (2021). 2% of those judges (notwithstanding both the unnamed judges and judges who ruled prior to 2021) were peremptorily disqualified. According to the California Commission on Judicial Performance’s 2021 Annual Report, a judge was disciplined for bias on 8 occasions: 5 times for “bias or appearance of bias not directed toward a particular class (includes embroilment, prejudgment, favoritism)” and 3 times for “bias or appearance of bias toward a particular class.”208Id. at 17. In 2021, three of the four private admonishments, id. at 40, and one of the eleven advisory letters dealt with bias, id. at 41–42. For context, there were “1,868 judgeships within the commission’s jurisdiction” including the judicial positions at the supreme court, courts of appeal, and superior courts.209Id. at 11. Even if all 8 instances of bias were from different superior court judges, less than 1% (0.5%) of all superior court judges would have faced discipline.

According to this data, there were more disqualified judges (2%) than judges disciplined for bias (0.5%). Judicial accountability was promoted when the 0.5% of judges who deviated from ethical guidelines were disqualified; what about the remaining 1.5% of judges? Of course, these judges might have just luckily evaded discipline for their bias. Discipline, unlike peremptory challenges, requires an investigation, not solely a mere allegation.210Id. at 10 (“[T]he standard of proof in [commission proceedings is] proof by clear and convincing evidence sufficient to sustain a charge to a reasonable certainty.” (citing Geiler v. Comm’n on Jud. Qualifications, 515 P.2d 1, 4 (Cal. 1973))). That being said, if the litigants and their attorneys truly thought their judge was biased, they could have complained to the Commission on Judicial Performance (at least anonymously) in order to avoid facing the same judge again.211Id. at 1.

A disqualified judge’s age, race, and gender, among other characteristics, were not easily identifiable. However, the judge’s political leanings (determined by which political party they were registered for) were discoverable for 21 out of the 37 disqualified judges. Thirteen judges (62%)  were Democrats, 7 judges (33%) were Republicans, and 1 judge (5%) was a Libertarian. This Note is committed to preserving these judges’ anonymity as they may understandably want to keep their politics confidential. Unfortunately, the distribution of party affiliation in the state judiciary was not readily ascertainable, but the total voter registration by political party provided some context—in 2021, 46.5% of voters were Democrats, and 24% of voters were Republican.21215-Day Report of Registration, Cal. Sec’y of State (Aug. 30, 2021), https://elections.cdn.sos.ca.gov/ror/15day-recall-2021/historical-reg-stats.pdf [https://perma.cc/2BTT-MGFN]. The comparison is illustrated by Figure 3 below:

Figure 3.

The law does not and cannot cover every kind of situation, so judicial discretion in the interpretation of the law maintains the legal system. Accordingly, judicial disqualification must take into account the diversity of experiences and legal philosophies that make up the bench. Each judge should have opportunities to arbitrate cases; otherwise, caselaw will cease to think outside the box. Yet, the data reveals that 62% of the disqualified judges were registered Democrats, and 33% of those judges were registered Republicans. Granted, without knowing how many California judges in sum have a Democratic-party affiliation, this is weak evidence for party-affiliation bias. But it is at least some insight that may suggest at best discriminatory effects and at worst purposeful discrimination against Democrat judges—for not only criminal, but also civil cases.213See infra Section II.C. If judges from a particular political party are systematically taken off matters through peremptory challenges, the judiciary becomes less representative. Consider how there are “persuasive correlations between the political party of the appointing authority and the judge’s decisions on certain issues,” according to an academic study of judicial decision-making.214Levi, supra note 23. Other group affiliations are also implicated: Black, Hispanic, and Asian voters are typically more liberal than conservative.215Midterm Election Preferences, Voter Engagement, Views of Campaign Issues, Pew Rsch. Ctr. (Aug. 23, 2022), https://www.pewresearch.org/politics/2022/08/23/midterm-election-preferences-voter-engagement-views-of-campaign-issues [https://perma.cc/U7N6-N5XD].

Of 37 disqualified judges, 33 of them had reviews on The Robing Room—a forum “by attorneys for attorneys” in which “judges are judged.”216FAQs, The Robing Room, http://www.therobingroom.com/california/FAQs.aspx?state=CA [https://perma.cc/FSY6-LW5S]. The Robing Room is “owned and operated by North Law Publishers, Inc., a New York Corporation, whose principal shareholders are attorneys.” Id. Fifteen judicial profiles had at least 1 comment from or before 2021 that mentioned a Section 170.6 motion—all but 1 recommended a peremptory challenge. Out of the 32 comments urging others to use Section 170.6 (many of which listed more than one reason), only 15 comments (47%)  complained of the judge’s bias. While 3 comments (9%) gave no reason at all, the remaining 17 comments cited explanations that did not concern bias: 17 (53%) for incompetence, 9 (28%) for unpleasant temperament, 4 (12.5%) for unnecessary delay, and 4 (12.5%) for disliked persons working in the judge’s chambers. Out of respect for these judges, their identities will remain anonymous, especially since the information is not necessary to this Note’s analytical aims. Figure 4 below demonstrates this distribution of motives:

Figure 4.

Section 170.6 blatantly spells out the one acceptable rationale for challenging the judge—bias. When litigants abuse their affidavit power against unbiased judges (that is, “judge shopping,” although the term does not quite capture the concept), they are admitting their search for a judge who will favor their side. The sample of comments from The Robing Room implies that Section 170.6 is not an obscure and hidden procedure. Rather, attorneys understand that they can peremptorily disqualify their judge through Section 170.6. Fifty-three percent of these comments recommending a challenge did not complain that the judge was biased. Admittedly, it is uncertain whether the litigants who challenged their judge in this study held the same beliefs as these reviewers or were influenced by these reviews in making their challenge. Though the data does not definitively prove that reasons outside of bias motivated these challenges, it still exposes what some practitioners think these challenges should be used for. Incompetence, unpleasant temperament, unnecessary delay, and disliked persons working in the judge’s chambers are undoubtedly serious problems, but they are problems that nevertheless affect both the plaintiff and defendant—there is no favoritism and therefore no bias. The duty of vigorous advocacy on behalf of the client is not a free pass for attorneys to bend the law to their will.

Besides The Robing Room, there were other published sources documenting criticism of the disqualified judges: circulating petitions for removal, judicial corruption activism pages, and news articles about their behavior in their private or public lives. One news outlet asked a disqualified judge about her alleged bias toward women to which she responded that there is probably an equally strong sentiment that she is biased toward men. Notably, the Commission on Judicial Performance publicly admonished one of the disqualified judges for improper conduct extraneous to bias.

There is arguably universal consensus that public confidence in the judiciary is of the utmost importance—the public needs assurance that they can rely on the courts for remedies to their legal grievances. The uproar against judges on the Internet (that is, the petitions, activism pages, news reports, and so forth) feeds the public impression that judicial independence is forgotten and left behind on the courthouse’s steps. Even if a judge’s attitudes on contentious issues in the legal and political community escape the public eye, the seemingly innocuous knowledge of the judge’s political-party registration can speak volumes given modern political polarization.217See Levi, supra note 23 (“[F]or judges to consider or present themselves as of different political teams . . . and for the experience of parties and lawyers to see judges so arrayed, would be highly destructive of the reality and appearance of fair and impartial, non-partisan courts.”). Republican litigants confronting Democrat judges may believe the “politicians in black robes” will unequivocally rule left, and vice versa for Democrat litigants. The perception of bias in the courts is disconnected from whether bias is actually rampant among judges.

IV.  ALTERNATIVES TO CALIFORNIA’S JUDICIAL PEREMPTORY CHALLENGE

A.  Existing Alternative Procedures

Despite an overall low risk of abuse, since judicial peremptory challenges are seemingly infrequent, there is little need for the challenge at all, at least in its current form. The empirical findings cast alternative approaches in a new light. This Note focuses on three ideas for compromise: the panel-exclusion approach, the interlocutory appeal approach, and the independent judge approach.

1.  The Panel-Exclusion Approach

The panel-exclusion approach advocates the adoption of a procedure similar to that used in arbitration.218See, e.g., Lab. Arb. Rules r. 12 (Am. Arb. Ass’n 2019) (“If the parties have agreed that the arbitrators shall appoint the neutral arbitrator from the National Roster, the AAA shall furnish to the party-appointed arbitrators . . . a list selected from the National Roster, and the appointment of the neutral arbitrator shall be made as prescribed in that section.”). Litigants would anonymously exclude judges who are randomly placed on the case’s panel.219Miller, supra note 147, at 482–83. Unlike the challenge as it currently stands, court administrators would provide litigants with a “compilation of numerous exclusion decisions,” including rates, prior to any challenge, so litigants can make informed decisions without relying solely on “mistakes in individual cases.”220Id. at 483–84. Disclosure of campaign activities could also help. Serbulea, supra note 45, at 1145 (“It would be difficult and costly for litigants to discover relevant information, so judges could be required to have on file copies of their campaign statements, as well as information on their campaign finances.”). This is an especially fruitful modification considering the overwhelming amount of unsolicited opinions and false stories online. Judicial analytics not only puts judges on notice about their inappropriate behavior and thus provides opportunities to cure such behavior, but also wins trust from the public by prioritizing transparency and honesty.

2.  The Interlocutory Appeal Approach

A retired Associate Justice of the Arkansas Supreme Court admires Tennessee’s civil procedure in which “[t]he judge refusing to recuse, following a motion to do so accompanied by an affidavit, must enter an order stating his or her reasons for not recusing and any other pertinent information from the record for an immediate, interlocutory appeal to the Tennessee Court of Appeals, where that court will expedite and conduct a de novo review.”221Justice Robert L. Brown, Retired, Judicial Recusal: It’s Time to Take Another Look Post-Caperton, 38 U. Ark. Little Rock L. Rev. 63, 73 (2015); see Tenn. Sup. Ct. R. 10B, § 2.01. Through the appeal, parties who failed to disqualify their judge would not have to endure a lengthy trial with an offended judge. Therefore, the challenge’s opponents might appreciate this third type of recourse before appeal of the entire case (joining review by a different judge, like the district’s chief judge, and mandamus review). The retired Associate Justice praises its efficacy in guarding judicial integrity and due process and urges Arkansas, a state where judges have discretion to deny disqualification motions without stating reasons, to follow suit.222Brown, supra note 221, at 73. His argument has merit in other jurisdictions with automatic disqualification, such as California, because appellate review will presumably lead to fewer judicial removals and prevent the public from falsely believing that there are more biased judges than is actually the case. The remedy provides no relief to litigants and attorneys who fear the insulted judge’s retaliation, but it should curtail judge shopping, especially on discriminatory grounds like race or gender, and minimize the odds of judicial intimidation. There is instinctive apprehension about crowding the appellate dockets, including the Supreme Court, but the Tennessee Administrative Office of the Courts—finding only ten or less appeals per year in a span of three years—and a Tennessee Court of Appeals judge—a “self-described ‘fan of the rule’ ”—calm these concerns.223Id. at 73 & nn.75–77. However, litigants lacking sufficient resources may not appeal even if their judge’s personal views have irreparably infected the proceeding.224Frost, supra note 40, at 571–72; see also Serbulea, supra note 45, at 1143 (“[F]inding an impartial appellate judge for an interlocutory appeal places a heavy burden on litigants.”).

3.  The Independent Judge Approach

Although his analysis revolves around a federal recusal statute’s reform, one legal scholar contributes a slightly different antidote to this dialogue. Another disinterested trial judge (that is, not the affected judge with a personal stake in the challenge) should rule on the disqualification motion because even the best-intentioned judge might be oblivious to their own faults. He further diverges from the affidavit procedure by suggesting that “the challenged judge be encouraged to file evidence refuting facts asserted in the recusal motion, and perhaps also an explanation of why disqualification is not justified,” so there is an “adversarial presentation of the issue.”225Frost, supra note 40, at 588; see, e.g., Tony Mauro, Courtside: When Planets Collide, Legal Times, Mar. 29, 2004, at 10 (“We are the only branch of government that must give reasons for what we do.”) (quoting Justice Kennedy); Serbulea, supra note 45, at 1142–43 (“It is . . . the judge who plays the role of the adversary party, but in an unfair way: getting to decide the matter, and rarely giving a reasoned (and written) explanation . . . . [J]udge impartiality[] is not consistent with the self-judging of recusal motions, which is the law in most states and the federal system . . . .”). Judges may worry about offending their colleagues,226Frost, supra note 40, at 552 (“Judges who wish to maintain collegial relations with one another hesitate to set in stone recusal procedures that might be viewed as disrespectful of their fellow judges.”). but given dissenting opinions and reversals of lower courts’ judgments, they are likely already accustomed to internal disagreement and can consequently stomach any potential discomfort.227See, e.g., United States v. Microsoft Corp., 253 F.3d 34, 116 (D.C. Cir. 2001) (per curiam) (disqualifying a district court judge from a highly publicized case even though the circuit judges who made the decision worked in the same courthouse as the disqualified judge). Additionally, he contends that “the appearance of justice will be better served, even if the actual rate of recusal remains unchanged.”228Frost, supra note 40, at 586.

B.  The Proposed Alternative Procedure

A more promising solution is a hybrid model between the panel-exclusion approach (specifically, the dissemination of judicial analytics) and the independent judge approach. After litigants receive the exclusion decisions and rates, they can file the motion with a different trial judge who will review both the motion and the challenged judge’s evidentiary explanation for factual and legal sufficiency. An interlocutory appeal is rendered unnecessary if an independent judge can accurately filter for allegations that actually deserve a judicial peremptory challenge. Admittedly, like federal law, this is not peremptory per se, but it will hopefully further reduce the number of “uninformed, misinformed, [and] delusional”229Raymond J. McKoski, Rewriting Judicial Recusal Rules with Big Data, 2020 Utah L. Rev. 383, 404. litigants who exercise the challenge.

As judges lack crisis managers, it is imperative that resources are invested into educating the public on judicial duties and verified statistics on judicial bias. Unlike news outlets, petitions, social media posts, and activism pages probably do not check the accuracy of the information they release. With the help of Big Data230See generally id. and artificial intelligence, judicial analytics will become more accurate over time, which will better alert litigants about the judges who are actually biased than other unverified sources. By making such information accessible, perhaps litigants will place less weight on factors like the judge’s race and gender, for example. The Robing Room states that slanderous comments posted in bad faith are subject to removal;231FAQs, supra note 216 (“We reserve the right to delete comments and ratings which we believe are libelous or not submitted in good faith.”). realistically, it can hardly stop all “sour grapes” who harbor disdain toward their judge for merely siding with the opposing party after fairly applying the law to the facts. Consider how one of the disqualified judges in the study asserted that there are an equal number of people who think she is biased toward either women or men. There is a possibility that the other sources noted in the study, besides the public admonishment, are campaigns by losing litigants to unjustifiably vilify their judge.

It seems problematic to allow judges to entertain motions petitioning their own disqualification, and the public agrees.232Press Release, Justice at Stake, Harris Interactive Public Opinion Poll on Judges and Money 1–2 (Feb. 12–15, 2009), https://www.brennancenter.org/sites/default/files/2009%20Harris%20Interactive%20National%20Public%20Opinion%20Poll%20on%20Judges%20and%20Money_0.pdf [https://perma.cc/MAD4-HQA4] (reporting that 81% of the surveyed public stated that judges should not decide motions calling for their recusal). An independent adjudicator should handle the challenge, so the challenged judge does not have to awkwardly decide their own neutrality. It is a win-win situation—if the reviewing judge finds the challenged judge corrupted with partiality, then the public will trust that the judiciary is void of collusion, but if the reviewing judge deems the challenged judge unbiased, then the public will have faith in the judiciary’s independence, and the judge will be guarded from discrimination. Some argue that the challenged judge is ideal because they are closest to the alleged facts;233Serbulea, supra note 45, at 1146 n.346. this is precisely the issue. That said, the reviewing judge may have a connection to the challenged judge—for example, a friendship—that skews their decision in favor of a denial. Perhaps the reviewing judge should show that there is no personal relationship to the challenged judge, but at some point, judges have to be trusted to rule fairly.

In lieu of an automatic transfer, judges can defend their fitness to serve and expose the litigants who use bias as a pretense for prohibited reasons, like discrimination based on party affiliation. It is easy for Section 170.6 to act as a Trojan horse carrying ulterior motives because automatic reassignment is swiftly delivered following a quick evaluation of procedural adequacy. If judges have no chance to prove the allegations of bias wrong, these litigants unwittingly trigger an endless feedback loop in which their baseless challenges inflate the number of “biased” judges which, in turn, instigates more challenges. The expectation is that fewer than 1.5% of judges will encounter peremptory disqualification, closing the disparity between the number of disqualified judges and the number of judges disciplined for bias. This will then demonstrate to the public that judges are not always predisposed to bias. To play devil’s advocate, the public may feel disheartened upon witnessing too many judges found guilty of bias, but the judiciary should commit to their integrity and weed out the “10% problem,”234Menell & Vacca, supra note 140, at 884–85. the estimate of incompetent judges. Judges inevitably do not command from Olympian heights; rather, they are subject to their beliefs and attitudes. It is difficult to procure solid evidence of judges’ unconscious biases,235See, e.g., Deborah Goldberg, James Sample & David E. Pozen, The Best Defense: Why Elected Courts Should Lead Recusal Reform, 46 Washburn L.J. 503, 525 (2007) (reporting that most people underestimate and undercorrect for their biases, according to social psychology research); Tobin A. Sparling, Keeping Up Appearances: The Constitutionality of the Model Code of Judicial Conduct’s Prohibition of Extrajudicial Speech Creating the Appearance of Bias, 19 Geo. L. Legal Ethics 441, 480 (2006) (“[J]udges may convince themselves they can rule fairly, unaware that the currents of bias often run deep.”). but litigants should still explain what manifestations, whether inside or outside the courtroom, caused them to suspect such biases.236For example, the judge addresses male attorneys as “counsel” but refers to female attorneys by their first name. They should not have free reign to judge shop due to a random feeling, especially in the U.S. legal system that is rooted in proof. Judicial economy is indeed lost if an inquiry is made into the merits as well,237Serbulea, supra note 45, at 1146 n.346. but it is a wiser alternative than the current framework that passively allows litigants to chase unequal justice.

V.  RECOMMENDATIONS FOR FUTURE RESEARCH

Regrettably, without formal statistical methods to control for irrelevant factors, this Note is unable to confirm a causal relationship between a judge’s peremptory disqualification and any complaints of bias on their Robing Room profile. For the same reason, it is unknown if failed judicial peremptory challenges affected case outcomes (comparing the case in which the challenge was initiated to any future cases with the disqualified judge—this would have informed the policy debate on whether disqualified judges are truly hostile). Therefore, the first recommendation for future research is setting up a more sophisticated study in order to discover results beyond mere correlations. Another recommendation is conducting surveys directed to (1) the public asking about their impression of judicial bias,238Surveys should be carefully formulated as people do not always answer honestly. See Park, supra note 121, at 571. (2) judges asking about their various group affiliations (especially characteristics that are not available through public materials such as race and ethnicity, gender, and age), and (3) attorneys asking what resources they have at their disposal when deciding whether they should use Section 170.6. Since judicial disciplinary proceedings for bias may miss judges with more subtle manifestations of bias, one suggestion is to conduct an experiment239See id. at 571–72 for one of the methods of uncovering unconscious biases. testing how widespread conscious and unconscious biases are among superior court judges in California. Ideally, this Note would analyze the relationship between the type of case (for example, personal injury) and judge shopping; unfortunately, there were not enough free and accessible documents online. For this pursuit, as a judge’s area of professional expertise is easy to find, future researchers should investigate whether movants of a specific type of case are strategically challenging judges with or without experience in that practice area. Lastly, the challenge’s prevalence is a regional phenomenon in the United States, so studies conducted in other states are recommended as well.

CONCLUSION

Notwithstanding limitations, the empirical data reported in this Note has value—it discovered that judicial peremptory challenges were quite rare and therefore abuse from these challenges was not out of hand. Among the few filed motions, most were automatically granted, indicating that the procedural protections were a flimsier shield than the statute had planned. Juxtaposing the higher percentage of disqualified judges with the lower percentage of judges reprimanded for bias implies that litigants are alleging bias as a mere formality. This is further corroborated by the finding that more than half of the comments on The Robing Room recommending others to challenge a certain judge did not mention bias.240To reiterate, this Note acknowledges issues other than bias—the California Legislature can determine whether these additional grounds for disqualification are warranted. In regard to discrimination, it found significantly more Democrat judges disqualified than Republican judges. Without additional research, this Note can only surmise about possible fixes to prevent discrimination, like the standard for employment law in which the judge’s ability to perform their duties must relate to the reasons for exclusion.241For an article proposing peremptory juror challenges to adopt this standard, see Ted A. Donner, Illinois Courts Struggle with Implicit Bias and Justice Stevens’s Legacy: Why Illinois Should Revisit His Dissenting Opinion in Purkett v. Elem, 53 Loy. U. Chi. L.J. 717, 745 (2022). Whether the challenge can inherit AB 3070 (the 2020 law that “requir[es] an attorney exercising peremptory strikes to show clear and convincing evidence [under an objective standard] that [their] action is unrelated to that juror’s membership in a protected group or class”242Gravdal, supra note 123 (emphasis omitted).) such that it is workable to judges poses an interesting question. Altogether, this Note concludes that there is not a serious risk of abuse from the challenge but Section 170.6 is still not a satisfactory remedy by legislation—there is no need to settle for less when there is a better solution. As Justice Kennedy wrote, judicial disqualification standards should extend beyond the minimum requirement of due process; however, they should not stretch so thin when judicial integrity is not completely broken. The proposed alternative will heal the issues produced when the challenge is granted as a matter of right by implementing an audit into the accusation’s truth. In other words, as a middle ground in the policy dichotomy, it will perfect the peremptory challenge and diminish the risk of abuse even more than the current model.

97 S. Cal. L. Rev. 253

Download

* Executive Senior Editor, Southern California Law Review, Volume 97; J.D. Candidate 2024, University of Southern California Gould School of Law; B.A. Communication 2019, University of California, Santa Barbara. Thank you to Professor Jonathan Barnett and Professor Robin Craig for serving as my advisors. To my family and friends—law school, much less this Note, would not be possible without your continuous support. I would also like to express my gratitude to the dedicated members of the Southern California Law Review for their hard work.

Restraining the Second Amendment in the Era of the Individual Right: Adopting a Modified South African Gun Control Model

In New York State Rifle & Pistol Association v. Bruen, the Supreme Court announced a novel historical test for judging the constitutionality of firearm laws. In combination with its earlier decisions in District of Columbia v. Heller and McDonald v. City of Chicago, the Court has created an onerous burden on federal and state legislatures attempting to regulate civilian firearm ownership. Given Heller’s individual right ruling, McDonald’s incorporation, and Bruen’s historical precedent requirement, it is clear that designing a restrictive firearm ownership system based on models that have proven successful in other Western countries is not possible, as most, if not all, of these would run afoul of these precedents. South Africa’s firearm licensing system, on the other hand, can provide a useful starting point for creating a framework that states can adopt. South Africa has significant private firearm ownership, its licensing system is not unduly restrictive, and it has proven successful in reducing gun violence. This Note therefore proposes adopting a version of South Africa’s firearm licensing system modified to survive judicial review in the United States. This Model Act likely represents close to the most restrictive licensing system that can pass judicial review following Bruen and might prove similarly effective in reducing gun violence in the United States.

There is almost no political question in the United States that is not resolved sooner or later into a judicial question.

—Alexis de Tocqueville1Alexis de Tocqueville, Democracy in America 257 (Harvey C. Mansfield & Delba Winthrop eds. & trans., 2000) (1835).

INTRODUCTION

The United States is in many ways an odd country, and there are few things more quintessentially American than the sheer quantity of firearms and relative frequency of mass shootings in this country. Presently, the United States has about 120 guns for every 100 citizens,2Global Firearms Holdings, Small Arms Surv. (Mar. 29, 2020), https://www.smallarmssurvey.org/database/global-firearms-holdings [https://perma.cc/6UY8-GCHP]. and a higher rate of gun violence than any other wealthy, developed country.3Gun Violence in the US Far Exceeds Levels in Other Rich Nations, Bloomberg (May 26, 2022, 5:00 PM), https://www.bloomberg.com/graphics/2022-us-gun-violence-world-comparison [https://perma.cc/TQF9-U8TY]. In addition to issues of gun violence, a high proportion of firearm ownership is closely associated with firearm suicide rates. See Michael Siegel & Emily F. Rothman, Firearm Ownership and Suicide Rates Among US Men and Women, 1981–2013, 106 Am. J. Pub. Health 1316, 1319 (2016) (finding a correlation between state-level firearm ownership and suicide rates of 0.71 among men and 0.49 among women). The tragic reality is that the gun control debate in the United States is never untimely. In light of the level of gun violence and ready availability of firearms in the United States, one solution seems simple: restrict access to firearms. After all, there is evidence that this approach can be successful.4See S. Chapman, P. Alpers, K. Agho & M. Jones, Australia’s 1996 Gun Law Reforms: Faster Falls in Firearm Deaths, Firearm Suicides, and a Decade Without Mass Shootings, 12 Inj. Prevention 365, 366 (2006). However, since the Second Amendment5U.S. Const. amend. II. has been interpreted to protect a broad, individual right to keep and bear arms,6District of Columbia v. Heller, 554 U.S. 570, 592 (2008); N.Y. State Rifle & Pistol Ass’n v. Bruen, 142 S. Ct. 2111, 2121 (2022). and any realistic prediction of the foreseeable future provides little reason to expect that the Second Amendment will be repealed, designing a perfect gun control statute from scratch is simply not an option. Moreover, our federal governmental structure further restricts our options. Promulgating a comprehensive, federal regulatory scheme that does not run afoul of the individual rights in, or the structural aspects of, the Constitution is infeasible.

Apart from the legal constraints on gun control, political, cultural, and economic roadblocks abound. Firearms are a prominent aspect of modern American culture,7Michael Waldman, The Second Amendment: A Biography 166 (2014); B. Bruce-Briggs, The Great American Gun War, 45 Pub. Int. 37, 41 (1976). and many Americans enjoy firearm ownership for safe, legitimate purposes like self-defense, hunting, and sport shooting.8Lydia Saad, What Percentage of Americans Own Guns?, Gallup (Nov. 13, 2020), https://news.gallup.com/poll/264932/percentage-americans-own-guns.aspx [https://perma.cc/B27Z-KY9R] (“Thirty-two percent of U.S. adults say they personally own a gun, while a larger percentage, 44%, report living in a gun household.”). Additionally, the United States firearms market is a $28 billion industry with significant lobbying strength.9Elizabeth MacBride, America’s Gun Business Is $28B. The Gun Violence Business Is Bigger, Forbes (Nov. 25, 2018, 5:00 AM), https://www.forbes.com/sites/elizabethmacbride/2018/11/25/americas-gun-business-is-28b-the-gun-violence-business-is-bigger [https://perma.cc/B2L2-UWUS]. Suffice it to say, there are significant constraints within which any regulatory structure must fit. Fortunately, however, it is not necessary to start from scratch. Foreign practices that have proven effective can be tailored to our constitutional constraints to develop a model gun control act for the several states to adopt. In particular, this Note proposes a distinctive approach: start with South Africa’s Firearms Control Act of 200010Firearms Control Act 60 of 2000 JSRSA (S. Afr.) (updated through 2014). and modify it to create an act that satisfies the U.S. Constitution and serves the policy goals of reducing access to firearms by those who would misuse them while keeping them available to responsible citizens. The purpose of this act is to provide a framework for the states to create a comprehensive firearm licensing system that can survive judicial review under current Second Amendment doctrine.11This Note proposes a model act for the states to adopt—rather than a proposed federal statute—because the amount of state and local law enforcement cooperation that would have to be demanded would be at risk of violating the anti-commandeering doctrine. See Printz v. United States, 521 U.S. 898, 929 (1997). This Model Act serves as a starting point for responsible firearm regulation in the era of the individual right and does not purport to be the final word on the Second Amendment question.

South Africa’s gun control system may be a surprising basis for a new federal gun control law in the United States; its intentional homicide rate far surpasses ours.12See Victims of Intentional Homicide, U.N. Office on Drugs and Crime https://dataunodc.un.org/dp-intentional-homicide-victims [https://perma.cc/2T6S-7XJQ]. However, South Africa has seen a steady decrease in gunshot-related deaths since it adopted the Firearms Control Act of 2000.13R. Matzopoulos, P. Groenewald, N. Abrahams & D. Bradshaw, Where Have All the Gun Deaths Gone?, 106 S. Afr. Med. J. 589, 590 (2016). The United States could see a reduction in its gunshot-related deaths by adopting a similar model. In any case, the large-scale empirical questions over the efficacy of various gun control systems are beyond the scope of this Note. The Note instead focuses on how we might adapt a comprehensive, firearms licensing scheme to our constitutional framework. South Africa’s model is an excellent starting point because it restricts access to especially dangerous firearms while providing individuals the opportunity to own firearms for self-defense, which the U.S. Supreme Court has said is the core right of the Second Amendment.14See District of Columbia v. Heller, 554 U.S. 570, 592 (2008). The Firearms Control Act of 2000 also does not completely prohibit ownership of AR-15’s and other similar firearms. This is important because a law completely banning AR-15’s and similar rifles could be in danger of being declared unconstitutional and setting an even more cumbersome precedent.15See Miller v. Bonta, No. 19-cv-01537, 2023 U.S. Dist. LEXIS 188421, at *97 (S.D. Cal. Oct. 19, 2023) (declaring California’s assault weapons ban unconstitutional). Additionally, other potential solutions devised to completely side-step the Supreme Court’s latest precedents are not only unlikely to succeed beyond perhaps the short term, but, if successful, could also create a worrying trend whereby state governments could close off its courts to citizens seeking to vindicate their constitutional rights. California, for instance, has created a one-way fee-shifting penalty that allows government defendants to recover costs from a plaintiff who loses on any claim in a case challenging a state or local firearm regulation, but never allows a plaintiff to recover attorneys’ fees from the government, even if the plaintiff wins on every claim.16Act of July 22, 2022, ch. 146, 2022 Cal. Stat. 15 (codified at Cal. Civ. Proc. Code § 1021.11(a) (West 2022)); see also Complaint for Declaratory, Injunctive, or Other Relief at 1, Miller v. Bonta, 646 F. Supp. 3d 1218 (S.D. Cal. 2022). If held constitutional,17At present, a federal district court has enjoined enforcement of this fee-shifting statute. Miller v. Bonta, 646 F. Supp. 3d 1218, 1227 (S.D. Cal. 2022); S. Bay Rod & Gun Club, Inc. v. Bonta, 646 F. Supp. 3d 1232, 1235 (S.D. Cal. 2022). this fee-shifting statute would chill future lawsuits by citizens seeking enforcement of their right to bear arms and would, at the very least, force them into a federal forum, unduly burdening the district courts.18If this practice of closing off state courts to claims the legislature does not want them to hear becomes widespread, federal courts would be unduly burdened with 42 U.S.C. § 1983 claims for rights the states do not want to respect. This has implications far beyond the gun control debate and could threaten other enumerated constitutional rights.19See Miller v. Bonta, 646 F. Supp. 3d 1218, 1224 (S.D. Cal. 2022) (“The principal defect of § 1021.11 is that it threatens to financially punish plaintiffs and their attorneys who seek judicial review of laws impinging on federal constitutional rights. Today, it applies to Second Amendment rights. Tomorrow, with a slight amendment, it could be any other constitutional right . . . .”) (footnotes omitted). Such jerry-rigging of procedural laws bearing on a constitutional right is bad policy that could encourage other states to similarly attempt to sabotage any constitutional right it wishes to infringe.20See Whole Woman’s Health v. Jackson, 595 U.S. 30, 65 (2021) (Sotomayor, J., concurring in part) (“[S]tate courts cannot restrict constitutional rights or defenses that our precedents recognize . . . . Such actions would violate a state officer’s oath to the Constitution.”). Rather than venturing down this destructive path, it is more effective to work within governing caselaw to achieve legitimate policy goals like gun safety and gun violence prevention.21A more dangerous exercise in legislative draftsmanship is to enact new statutes that criminalize ownership of commonly owned weapons like the AR-15. See, e.g., 720 Ill. Comp. Stat. 5/24–1.9(b) (2023). Acts such as these are not only unlikely to survive judicial review but could also create sweeping precedent severely limiting how a future Supreme Court might approach the Second Amendment question. In fact, within days, three lawsuits were filed in federal and state court, challenging the law as unconstitutional. See Mitch Smith, Illinois Passed a Sweeping Ban on High-Powered Guns. Now Come the Lawsuits., N.Y. Times (Jan. 20, 2023), https://www.nytimes.com/2023/01/20/us/illinois-gun-ban-second-amendment.html [https://perma.cc/X7P7-4CKV]. Even if many or all of the Supreme Court’s recent Second Amendment cases were incorrectly decided, they remain binding precedent.22As any realist would point out, we have no choice but to adhere to precedent in the Second Amendment context at least until the composition of the Supreme Court changes dramatically. At minimum, two of the six conservative Supreme Court Justices (Chief Justice Roberts and Justices Thomas, Alito, Gorsuch, Kavanaugh, and Barrett) would have to be replaced to provide a chance to overrule District of Columbia v. Heller, 554 U.S. 570 (2008), McDonald v. City of Chicago, 561 U.S. 742 (2010), or N.Y. State Rifle & Pistol Association v. Bruen, 142 S. Ct. 2111 (2022). Thus, a detailed look at the Court’s recent Second Amendment precedents is necessary to develop a statutory scheme that will survive judicial review.

Part I begins with a breakdown of the Supreme Court’s recent Second Amendment jurisprudence from Heller through Bruen, analyzing the various doctrines articulated in these cases. Part I ends with a summary of the constitutional limitations that the model gun control statute must satisfy. Part II summarizes the salient points of South Africa’s Firearms Control Act of 2000. Part III provides the full text of the Model Firearms Control Act. Part IV argues that this Model Act is likely to be upheld by the Court.

I.  CONSTITUTIONAL LIMITATIONS OF GUN CONTROL

Second Amendment jurisprudence has been rather scant from its ratification in 1791 until the Heller decision in 2008, when the Amendment took on its modern meaning.23See generally District of Columbia v. Heller, 554 U.S. 570 (2008). Before the swell in revisionist legal scholarship that began in the 1960s, “[t]here was no more settled view in constitutional law than that the Second Amendment did not protect an individual right to own a gun.”24Waldman, supra note 7, at 97. Yet, the Second Amendment is now interpreted to protect a broad, individual right to own a firearm for self-defense.25See Heller, 554 U.S. 570; McDonald, 561 U.S. 742; Bruen, 142 S. Ct. 2111. Because of this recent, dramatic shift in case law, an investigation into the Court’s modern Second Amendment jurisprudence is required to create a model gun control act that is likely to survive judicial review.

A.  The Trilogy of Modern Second Amendment Jurisprudence

In a trilogy of cases—District of Columbia v. Heller,26Heller, 554 U.S. at 570. McDonald v. City of Chicago,27McDonald, 561 U.S. at 742. and New York State Rifle & Pistol Association v. Bruen28Bruen, 142 S. Ct. at 2121.—the Supreme Court established a broad, individual right to keep and bear arms, irrespective of any militia service, effective against both the federal and state governments. This broad protection of firearm ownership is built on four main principles: (1) the individual right approach, (2) the pre-existing right doctrine, (3) the common use doctrine, and (4) incorporation. To shape a gun control scheme to fit within controlling Supreme Court precedent, these four doctrines flowing from this trilogy of cases must be mapped out and understood.

1.  The Necessity of an Individual Right

Before the Court would have an opportunity to incorporate the Second Amendment to the states, it had to lay some precedential groundwork to convert the Second Amendment into a right that could be incorporated. To a large extent, finding an individual right to keep and bear arms for the purpose of self-defense without any militia service requirement was a prerequisite to incorporating the right. Prior to Heller, scholars and jurists had proposed three main approaches to interpreting the Second Amendment.29See David A. Lieber, Comment, The Cruikshank Redemption: The Enduring Rationale for Excluding the Second Amendment from the Court’s Modern Incorporation Doctrine, 95 J. Crim. L. & Criminology 1079, 1080–81 (2005). First, the “collective right” approach argued that the right to keep and bear arms protected the right of the states to arm and organize militias.30Id. at 1080. Second, the “limited individual right” or “sophisticated individual right” approach suggested that the right to keep and bear arms does protect an individual right, but only to the extent that individuals participate in a well-regulated militia.31Id. at 1080–81. Third, the unmodified “individual right” approach embraced the idea that the Second Amendment protects an individual’s right to keep and bear arms irrespective of any participation in a well-regulated militia, essentially reading the prefatory clause out of the amendment.32Id. at 1081.

Throughout pre-Heller Second Amendment case law and scholarship, the individual right approach was overwhelmingly disfavored.33From 1888, when law review articles began to be indexed, to 1960, no law review articles concluded that the Second Amendment guaranteed an individual right. Waldman, supra note 7, at 97. The first law review article to argue otherwise, published in 1960, was a student-written note which concluded that the Second Amendment provided a “right of revolution” that the Southern States availed themselves of during the Civil War. Stuart R. Hays, The Right to Bear Arms, A Study in Judicial Misinterpretation, 2 Wm. & Mary L. Rev. 381, 387–88 (1960). Between 1970 and 1989, however, twenty-five articles endorsing the individual right were written, at least sixteen of which were written by lawyers who had represented or been employed by the National Rifle Association (“NRA”) or other gun rights organizations. Carl T. Bogus, The History and Politics of Second Amendment Scholarship: A Primer, 76 Chi.-Kent L. Rev. 3, 8 (2000). From ratification until 2001, no federal appellate court had ever endorsed the individual right approach to the Second Amendment,34See Lieber, supra note 29, at 1097–98. and the first case to adopt this approach, United States v. Emerson,35United States v. Emerson, 270 F.3d 203 (5th Cir. 2001), abrogated by United States v. Rahimi, 61 F.4th 443 (5th Cir. 2023). did so only in dictum.36Id. at 260; Lieber, supra note 29, at 1081. Shortly following Emerson, then Attorney General John Ashcroft issued a memorandum to all United States Attorneys stating that the individual right approach reflects the correct understanding of the Second Amendment, reversing the Department of Justice’s longstanding policy regarding Second Amendment interpretation.37Memorandum from Attorney General John Ashcroft to All United States’ Attorneys (Nov. 9, 2001), https://www.justice.gov/archives/ag/attorney-general-memorandum-regarding-5th-circuit-united-states-court-appeals-decision-united [https://perma.cc/T6NT-XKH5]; Lieber, supra note 29, at 1081. Had the Court endorsed the collective right approach in Heller, as it had 132 years earlier,38United States v. Cruikshank, 92 U.S. 542, 549 (1875), overruled in part by McDonald v. City of Chicago, 561 U.S. 742 (2010). the right to bear arms would essentially be a right belonging to the states, making incorporation nonsensical as a state could not meaningfully infringe its own right.39Possession of a right implies the possession of an option. See Right, Black’s Law Dictionary (11th ed. 2019). It therefore follows that a decision to not exercise a right is unassailable. Thus, incorporating the “collective right” of a state to arm its own militias would be nonsensical, since it would have a concomitant right to not arm its militias. Alternatively, had the Court endorsed the limited individual right in Heller, a subsequent decision incorporating that right would only prevent states from disarming individuals serving in its own militias, which would provide no protection to anyone outside the National Guard.40See Lieber, supra note 29, at 1080–81, 1120. Thus, it was necessary for the Court to find an individual right to keep and bear arms in the Second Amendment, independent of any militia service, to meaningfully incorporate that amendment against the states.

2.  The Pre-Existence Doctrine: Finding the Individual Right in Text and History

In finding a free-standing, individual right to bear arms in the Second Amendment, the Court relied on the notion of some constitutional rights having pre-existed the ratification of the clauses protecting them.41District of Columbia v. Heller, 554 U.S. 570, 592 (2008) (“[T]he Second Amendment, like the First and Fourth Amendments, codified a pre-existing right.”). Although the Court did not cite any authority for this proposition, this quote from Heller has been parroted by numerous cases and law review articles, but there is a paucity of literature or case law substantively discussing the idea that the First and Fourth Amendments codified a pre-existing right. The discussion of the pre-existence and codification of the rights enshrined in the First and Fourth Amendments scarcely goes deeper than to quote Heller, and possibly to analogize the First Amendment to the Second. See, e.g., David B. Kopel, The First Amendment Guide to the Second Amendment, 81 Tenn. L. Rev. 417, 419 (2014) (“[T]he Supreme Court has strongly indicated that First Amendment tools should be employed to help resolve Second Amendment issues.”); Tyler v. Hillsdale Cnty. Sherriff’s Dep’t, 837 F.3d 678, 711 (6th Cir. 2016) (Sutton, J., concurring in part) (“The First Amendment offers a useful analogy [to the Second Amendment].”); United States v. Marzzarella, 614 F.3d 85, 96–97 (3d Cir. 2010) (applying a sliding scale test to the Second Amendment whereby the stringency of the standard varies according to the degree to which the statute burdens the right), abrogated by N.Y. State Rifle & Pistol Ass’n v. Bruen, 142 S. Ct. 2111 (2022). According to the Court, the Second Amendment did not create a new right but constitutionalized a pre-existing right.42Heller, 554 U.S. at 592 (“[T]his is not a right granted by the Constitution. Neither is it in any manner dependent on that instrument for its existence. The second amendment declares that it shall not be infringed . . . .”) (quoting United States v. Cruikshank, 92 U.S. 542, 553 (1876)). This pre-existence argument relies on the proposition that the framers of the Second Amendment intended to codify a right to bear arms that already existed in English law43See id. at 593–94. and simply wished to create a stronger protection for it. The Court purported to find a textual basis for this conclusion, stating that “[t]he very text of the Second Amendment implicitly recognizes the pre-existence of the right and declares only that it ‘shall not be infringed.’ ”44Id. at 592. Even if it is assumed that the text of the amendment implies a pre-existing right, it is not clear that this right comes from old English and Colonial law. An at least equally plausible explanation is that the Second Amendment confirms that the federal government does not have the power to disarm state militias. See The Federalist No. 46 (James Madison).

The Heller Court began its historical analysis by stating that “[t]he Second Amendment is naturally divided into two parts: its prefatory clause and its operative clause.”45Heller, 554 U.S. at 577. The prefatory clause states “[a] well regulated Militia, being necessary to the security of a free State . . . .” The operative clause states that “the right of the people to keep and bear arms shall not be infringed.”46Id. at 579–98; U.S. Const. amend. II. The Court asserted that the prefatory clause announces only the amendment’s justification, and does not limit the scope of the operative clause.47Heller, 554 U.S. at 577–78. After its explication, the Court concluded that the prefatory clause “fits perfectly” with an operative clause understood to grant an individual right to keep and bear arms because the pre-constitutional history showed that tyrants had eliminated militias not by banning them but by disarming them.48Id. at 598.

The Court supported its individual right approach through a sort of reverse incorporation argument limited to “analogous arms-bearing rights in state constitutions that preceded and immediately followed adoption of the Second Amendment.”49Id. at 600–01; see Joseph Blocher, Reverse Incorporation of State Constitutional Law, 84 S. Cal. L. Rev. 323, 381 (2011). Although many of the state constitutions had more individualistic wording,50Heller, 554 U.S. at 600–03. Even at the time Heller was being decided, the vast majority of states recognized an individual right to keep and bear arms. Eugene Volokh, State Constitutional Rights to Keep and Bear Arms, 11 Tex. Rev. L. & Pol. 191, 192 (2006) (concluding forty-four states recognize an individual right to bear arms); Adam Winkler, Scrutinizing the Second Amendment, 105 Mich. L. Rev. 683, 686, 711 (2007) (concluding that forty-two states protect an individual right to bear arms). the Court did not take this to conclude that the Second Amendment was materially different from its state analogues. To the contrary, the Court used the more individual rights-focused arms-bearing provisions of state constitutions—and state supreme court decisions interpreting those provisions—to read the Second Amendment as conferring a broad individual right.51Heller, 554 U.S. at 600–03. Pennsylvania’s Declaration of Rights of 1776 read “the people have a right to bear arms for the defence of themselves and the state . . . .”52Id. at 601; Pa. Const. of 1776, art. I, cl. 13, amended by Pa. Const. art. I, § 21 (emphasis added). and Vermont’s 1777 Declaration of Rights contained a nearly identical provision.53Heller, 554 U.S. at 601; Vt. Const. of 1777, ch. I, cl. XV, amended by Vt. Const. ch I, art. XVI. The Court further supported its argument by describing roughly contemporaneous state analogues.

Between 1789 and 1820, nine States adopted Second Amendment analogues. Four of them—Kentucky, Ohio, Indiana, and Missouri—referred to the right of the people to “bear arms in defence of themselves and the State.” Another three States—Mississippi, Connecticut, and Alabama—used the even more individualistic phrasing that each citizen has the “right to bear arms in defence of himself and the State.” Finally, two States—Tennessee and Maine—used the “common defence” language of Massachusetts.54Heller, 554 U.S. at 602–03 (citations omitted).

The Court noted that the decision of at least seven of these nine states to unequivocally protect an individual right to bear arms is strong evidence that the framers of the Second Amendment conceived of the right to bear arms as an individual right.55Id. at 603. Contrary to the Court’s conclusion, however, the inclusion of language clearly protecting an individual right to bear arms in state constitutional analogues to the Second Amendment might be indicative of a structural difference between state and federal governments. The federal right to bear arms could simply prevent the federal government from disarming state militias while states might be best understood to have the right to arm and disarm their own militias and citizens as they see fit. For further discussion on the incorporation issue, which is by its very nature intertwined with the individual right issue, see infra Section I.A.4.

The Court next sought precedential support for its individual right interpretation.56Id. at 600–01. The Court first cited Nunn v. State, an 1846 case in which the Georgia Supreme Court struck down a ban on carrying pistols openly, stating that the Second Amendment protects “the natural right of self-defense.”57Nunn v. State, 1 Ga. 243, 251 (1846). The Heller Court noted that the Georgia Supreme Court “perfectly captured the way in which the operative clause of the Second Amendment furthers the purpose announced in the prefatory clause, in continuity with the English right.”58Heller, 554 U.S. at 612. Despite what the Heller Court and Tennessee Supreme Court’s wording might suggest, it is important to note that the “English right” in question is not easily analogized to the Second Amendment. Importantly, the right to bear arms for self-defense in the pre-constitutional English right contains clearer limiting language and was a concession by the English Crown and subject to the will of parliament. See Bill of Rights 1689 1 W. & M., 2d sess. c. 2, § 7; see also 1 William Blackstone, Commentaries on the Laws of England *130 (William Carey Jones ed., Claitor’s Publ’g Div. 1976) (1765). In further support of its position, the Court cited State v. Chandler, an 1850 case in which the Louisiana Supreme Court held that United States constitution guaranteed citizens the right to carry arms openly.59State v. Chandler, 5 La. Ann. 489, 490 (1850). In response to the dissent’s reliance on Aymette v. State, an 1840 decision in which the Tennessee Supreme Court adopted a limited individual right approach for its own state constitutional right to bear arms,60Aymette v. State, 21 Tenn. 154, 161 (1840) (“[W]e must understand the expressions as . . . relating to public, and not private, to the common, and not the individual, defence.”). the Court reasoned that more important than this decision was the Tennessee Supreme Court’s later decision in Andrews v. State.61Andrews v. State, 50 Tenn. 165 (1871). In Andrews, the Tennessee Supreme Court concluded that its state constitutional right to bear arms protected the right to bear arms for personal self-defense, overruling Aymette.62Id. at 178–79. However, the relevant state constitutional provision reads: “[T]he citizens of this State have a right to keep and to bear arms for their common defense; but the Legislature shall have power, by law, to regulate the wearing of arms with a view to prevent crime.” Tenn. Const. art. I, § 26. This is notable because the text of the Tennessee Constitution’s arms-bearing provision is manifestly different from the text of the Second Amendment. The Andrews court itself held—on anti-incorporation grounds—that the Second Amendment does not protect a right to bear arms for self-defense against state infringement. Andrews, 50 Tenn. at 175, 178–79. In other words, Tennessee’s counterpart to the Second Amendment protected an individual right where the Second Amendment did not. This indicates that the Andrews court considered the Second Amendment to be not only meaningfully different from, but also narrower than, its state counterpart. Id.; see also Simpson v. State, 13 Tenn. 356, 360 (1833) (construing the state constitution to protect an individual right to bear arms); cf. State v. Reid, 1 Ala. 612, 616 (“[T]he act, ‘To suppress the evil practice of carrying weapons secretly,’ [does not] trench upon the [Alabama] constitutional rights of the citizen.”).

Turning to its own precedents, the Court asked whether any of its prior decisions foreclosed its ultimate conclusion in Heller. The Court began with its decision in United States v. Cruikshank,63United States v. Cruikshank, 92 U.S. 542 (1875), overruled in part by McDonald v. City of Chicago, 561 U.S. 742 (2010). in which the Court vacated a white mob’s convictions for depriving black militia men of their right to bear arms, holding that the Second Amendment “means no more than that it shall not be infringed by Congress.”64Id. at 553. The Heller Court reasoned that there was no claim in Cruikshank that the defendants had violated the victims’ right to carry arms in a militia, and that the Cruikshank Court’s discussion made little sense if it was speaking of a collective rather than an individual right.65District of Columbia v. Heller, 554 U.S. 570, 620 (2008). The Court rests this argument on the Cruikshank Court’s conclusion that “ ‘the people [must] look for their protection against any violation by their fellow-citizens of the rights it recognizes’ to the States’ police power.”66Id. (quoting Cruikshank, 92 U.S. at 553) (alteration in original).

The Heller Court next turned to United States v. Miller,67United States v. Miller, 307 U.S. 174 (1939), abrogated by McDonald v. City of Chicago, 561 U.S. 742 (2010). reasoning that it not only failed to foreclose the possibility of an individual right, but also “positively suggests” it.68Heller, 554 U.S. at 622. Miller considered whether a law prohibiting the unregistered possession of a short-barreled shotgun ran afoul of the Second Amendment.69Miller, 307 U.S. at 175–76. In concluding that it did not, the Miller Court announced its interpretation of the Second Amendment’s purpose.

In the absence of any evidence tending to show that possession or use of a ‘shotgun having a barrel of less than eighteen inches in length’ at this time has some reasonable relationship to the preservation or efficiency of a well regulated militia, we cannot say that the Second Amendment guarantees the right to keep and bear such an instrument.70Id. at 178.

The Court reasoned that the Miller Court’s basis for concluding that the Second Amendment did not apply was not that the Second Amendment failed to protect non-military use, but that it did not protect the type of firearm at issue.71Heller, 554 U.S. at 622. Before announcing its “common use” doctrine, however, the Court acknowledged some limitations on the individual right to bear arms, such as the historical precedence for prohibiting the public carry of “dangerous and unusual weapons.”72Id. at 627 (citing 4 William Blackstone, Commentaries on the Laws of England *149 (William Carey Jones ed., Claitor’s Publ’g Div. 1976) (1765) (“The offense of riding or going armed with dangerous or unusual weapons is a crime against the public peace . . . .”)).

3.  Market Share as Constitutionality: The Common Use Doctrine

With the individual right in hand, the Heller Court turned to the law at issue, which totally banned handgun possession in the home and required any lawfully owned firearm to be disassembled and bound by a trigger lock.73Id. at 628. In determining that the law was unconstitutional, the Court began by concluding that “the inherent right to self-defense has been central to the Second Amendment right.”74Id. True enough, there is no serious doubt that the right to self-defense predates the constitution as part of the common law75See, e.g., Blackstone, supra note 58. and continues to exist in the United States today.76See Restatement (Second) of Torts §§ 63–68 (Am. L. Inst. 1965). It would be a novel legal principle indeed to compel citizens to allow themselves to be victimized by an aggressor. The right to bear arms to effectuate this defense of life and limb also existed in England prior to the ratification of the U.S. Constitution, at least by statute, as a “public allowance, under due restrictions, of the natural right of resistance and self-preservation . . . .”77Blackstone, supra note 58, at *144; see Bill of Rights 1689 1 W. & M., 2d sess. c. 2, § 7. Although the statutory right to bear arms for self-defense in England was considered less fundamental than the right to self-defense in general,78Compare Blackstone, supra note 58 (“Both the life and limbs of a man are of such high value, in the estimation of law of England, that it pardons even homicide if committed se defendendo (in self-defense), or in order to preserve them.”), with Blackstone, supra note 58, at *144 (“The . . . last auxiliary right of the subject, that I shall at present mention, is that of having arms for their defense, suitable to their condition and degree, and such as are allowed by law.”) (emphasis added). and was a concession by the Crown that presupposed an omnipotent legislature—a feature clearly absent from our constitutional scheme—the Court has insisted on the centrality of individual self-defense to the right to bear arms.79See District of Columbia v. Heller, 554 U.S. 570, 628 (2008); N.Y. State Rifle & Pistol Ass’n v. Bruen, 142 S. Ct. 2111, 2125 (2022). There is, however, significant historical evidence to the contrary. See William Carey Jones, Annotation, Blackstone, supra note 58, at *144 n.20 (“The constitutional right to bear arms in this country does not mean the right to bear them for individual defense . . . .”); Andrews v. State, 50 Tenn. 165, 197 (1871); United States v. Cruikshank, 92 U.S. 542, 591–92 (1875), overruled in part by McDonald v. City of Chicago, 561 U.S. 742 (2010); The Federalist No. 46 (James Madison) (describing the rationale for the Second Amendment in terms of militia service); see also Waldman, supra note 7, at 6 (explaining that keeping arms for English militia service was not an individual right but a duty); see generally Saul Cornell & Nathan DeDino, A Well Regulated Right: The Early American Origins of Gun Control, 73 Fordham L. Rev. 487 (2004).

The purported centrality of self-defense to the Second Amendment, combined with the individual right approach, allowed the Court to announce a new, sweeping doctrine in Heller. The Court reasoned that “[u]nder any of the standards of scrutiny that we have applied to enumerated constitutional rights, banning from the home ‘the most preferred firearm in the nation to “keep” and use for protection of one’s home and family,’ would fail constitutional muster.”80Heller, 554 U.S. at 628–29 (quoting Parker v. District of Columbia, 478 F.3d 370, 400 (2007)); see also Gary Kleck & Marc Gertz, Armed Resistance to Crime: The Prevalence and Nature of Self-Defense with a Gun, 86 J. Crim. L. & Criminology 150, 182–83 (1995). It noted that few laws in our nation’s history have come close to the restriction the District of Columbia has imposed and several of those laws have been struck down.81Heller, 554 U.S. at 629. Because handguns have been overwhelmingly chosen by the American people as their preferred arm for self-defense, a complete prohibition of its use runs afoul of the individual right to bear arms for the very purpose of self-defense.82Id. This common use doctrine begs the question: If it is unconstitutional to outright ban firearms in common use for self-defense, how would the Court approach bans on classes of arms which are not in common use because they were banned before they could get into common use?83For instance, the National Firearms Act has capped the market of machine guns by only allowing the lawful possession and transfer of machine guns lawfully owned prior to May 19, 1986. 27 C.F.R. § 479.105(b) (2023). This imposed market cap means that machine guns no longer have the chance to get into common use. It is not clear whether the Heller decision means that such a law is unconstitutional. The Court did not address this question,84It is true, however, that if the Second Amendment was intended to protect an individual right to bear arms for the purpose of self-defense—as indeed the Court has held—there must be some allowance made for citizens to keep and bear modern weapons. If citizens could only keep and bear arms in use at the time the Amendment was ratified, the right would be meaningless today. noting that it did not undertake an analysis of the full scope of the Second Amendment.85Heller, 554 U.S. at 626–27. However, the Court stated that nothing in its opinion “should be taken to cast doubt on longstanding prohibitions on the possession of firearms by felons and the mentally ill, or laws forbidding the carrying of firearms in sensitive places such as schools and government buildings, or laws imposing conditions and qualifications on the commercial sale of arms.”86Id. In fact, the Court noted that the measures it listed are presumptively lawful and that its list was inexhaustive.87Id. at 627 n.26. This is an important concession by the Court because by noting that its list of presumptively lawful measures was inexhaustive, the Court indicated that it might be open to other presumptively lawful restrictions to the right to bear arms, so long as there is a historical precedent that is satisfactory in the Court’s view.

4.  Incorporation

The Supreme Court would of course go on to conclude in McDonald v. City of Chicago that the right to bear arms is “deeply rooted in this Nation’s history and tradition”88McDonald v. City of Chicago, 561 U.S. 742, 768 (2010) (quoting Washington v. Glucksberg, 521 U.S. 702, 721 (1997)). and incorporate the Second Amendment in full.89Id. at 791. In so doing, it relied heavily on Heller’s individual right approach and common use doctrine, arguing that history and precedent pointed “unmistakably” to the conclusion that the Second Amendment is “deeply rooted” in our “history and tradition.”90Id. at 767–70. Just as in Heller, the Court argued that the right to bear arms for self-defense was as fundamental as the broader right self-defense.91Id. at 768. Confusingly, the Court stated that “by 1765, Blackstone was able to assert that the right to keep and bear arms was ‘one of the fundamental rights of Englishmen.’ ” Id. (quoting Heller, 554 U.S. at 594). This is a quote from Heller, but not from Blackstone, who in fact listed the right to bear arms as an auxiliary right, not a fundamental one. See Blackstone, supra note 58, at *144 (“The . . . last auxiliary right of the subject, that I shall at present mention, is that of having arms for their defense . . . .”) (emphasis added). In incorporating the individual right to the states, the Court had another perfect occasion to utilize the doctrine of reverse incorporation92See Blocher, supra note 49. to adopt a standard of review based on how state supreme courts have analyzed their own constitutions’ arms-bearing provisions that the Court saw as analogous to the Second Amendment. Most states recognize an individual right to keep and bear arms but allow “reasonable regulations” restricting that right.93Id. at 383; Winkler, supra note 50, at 686–87. Despite the states’ far greater experience in drafting and reviewing gun laws, the Supreme Court left the decision over what standard applied to Second Amendment cases to another day, eventually settling on Bruen’s historical test.94N.Y. State Rifle & Pistol Ass’n v. Bruen, 142 S. Ct. 2111, 2129–30 (2022).

The confluence of the individual right approach, the common use doctrine, and incorporation has opened many long-standing state firearms laws to constitutional scrutiny, even before Bruen was decided. California, for instance, has prohibited the purchase, sale, and manufacture of high-capacity magazines95California defines high-capacity or “large capacity magazines” as “any ammunition feeding device with the capacity to accept more than 10 rounds . . . .” Cal. Penal Code § 16740 (West 2012). The terms “high-capacity magazine” and “large-capacity magazine” are used interchangeably in this Note. since 2000,96See Cal. Penal Code § 32310 (West 2012 & Supp. 2020). and by popular initiative in 2016 expanded the prohibition to make possession of high-capacity magazines a felony offense, regardless of the date the magazine was acquired.97Id.; Safety for All Act, 2016 Cal. Legis. Serv. Prop. 63, § 6.1 (West), adding Cal. Penal Code § 32310(c)–(d) (Supp. 2020). This new law gave rise to protracted but groundbreaking litigation. In Duncan v. Becerra,98Duncan v. Becerra, 366 F. Supp. 3d 1131 (S.D. Cal. 2019), rev’d sub nom. Duncan v. Bonta, 19 F.4th 1087 (9th Cir. 2021), vacated, 142 S. Ct. 2895 (2022) (mem.). the outright ban on possession of high-capacity magazines was ruled unconstitutional as a Fifth Amendment taking without just compensation and as violative of the Second Amendment because it imposed a substantial burden on the right to self-defense and the right to keep and bear arms.99Id. at 1185–86. The district court enjoined the statute, and its decision was affirmed on appeal by the Ninth Circuit,100Duncan v. Becerra, 970 F.3d 1133, 1141 (9th Cir. 2020), vacated sub nom. Duncan v. Bonta, 142 S. Ct. 2895 (2022). but was later reversed on rehearing en banc.101Duncan v. Bonta, 19 F.4th 1087, 1096 (9th Cir. 2021) (en banc), vacated, 142 S. Ct. 2895 (2022). The Supreme Court then vacated the judgement and remanded the case to the Ninth Circuit for further consideration in light of its decision in Bruen.102Duncan v. Bonta, 142 S. Ct. 2895, 2895 (2022). On remand from the Ninth Circuit, the District Court once again held California’s high-capacity magazine ban unconstitutional, but stayed its order enjoining enforcement while the California Attorney General appealed the decision.103Duncan v. Bonta, No. 17-cv-1017, 2023 U.S. Dist. LEXIS 169577 (S.D. Cal. Sept. 22, 2023), appeal docketed, No. 23-55805, 2023 U.S. App. LEXIS 25723 (9th Cir. Sept. 28, 2023). It therefore remains to be seen how the latest Supreme Court precedent will affect this high-capacity magazine ban, but it suffices to say that the law in this area remains very much in flux.

5.  The Third Act: Applying Heller to Public Carry Licensing

Building on the bedrock of the individual right principle, the common use doctrine, and the Second Amendment’s incorporation, the Court recently expanded the Amendment’s protections with its historical precedence doctrine. At issue in Bruen was a New York law that made it a crime to possess a firearm without a license.104N.Y. State Rifle & Pistol Ass’n v. Bruen, 142 S. Ct. 2111, 2122 (2022). New York’s provision for licenses to carry firearms outside the home for self-defense was particularly stringent. An applicant could not obtain that license without a showing of “proper cause.”105Id. at 2123 (citing N.Y. Penal Law. § 400.00(2)(f) (McKinney 2022)). Without this showing of “proper cause,” an applicant may only obtain a “restricted” license to carry a firearm for such purposes as “hunting, target shooting, or employment.” Id. New York courts have defined “proper cause” as requiring an applicant to “demonstrate a special need for self-protection distinguishable from that of the general community”106Klenosky v. N.Y.C. Police Dep’t, 428 N.Y.S.2d 256, 257 (N.Y. App. Div. 1980), abrogated by N.Y. State Rifle & Pistol Ass’n v. Bruen, 142 S. Ct. 2111 (2022). such as evidence “of particular threats, attacks or other extraordinary danger to personal safety.” Living or working in a high-crime area was considered insufficient to demonstrate proper cause.107See Bernstein v. Police Dep’t of N.Y.C., 445 N.Y.S.2d 716, 717 (N.Y. App. Div. 1981), abrogated by N.Y. State Rifle & Pistol Ass’n v. Bruen, 142 S. Ct. 2111 (2022).

To evaluate the constitutionality of the New York law, the Bruen Court began by clarifying the test for Second Amendment challenges. The Court noted that the circuit courts had coalesced around a two-part test that combined a historical inquiry with means-end scrutiny, but it rejected this approach.108Bruen, 142 S. Ct. at 2125–26. The Court instead leaned on its historical approach from Heller and specifically rejected any interest balancing test,109Id. at 2127 (“Heller and McDonald do not support applying means-end scrutiny in the Second Amendment context. Instead, the government must affirmatively prove that its firearms regulation is part of the historical tradition that delimits the outer bounds of the right to keep and bear arms.”); id. at 2131 (“The Second Amendment ‘is the very product of an interest balancing by the people’ and it ‘surely elevates above all other interests the right of law-abiding, responsible citizens to use arms’ for self-defense.”) (emphasis in original) (quoting District of Columbia v. Heller, 554 U.S. 570, 635 (2008)). settling on the following standard:

When the Second Amendment’s plain text covers an individual’s conduct, the Constitution presumptively protects that conduct. The government must then justify its regulation by demonstrating that it is consistent with the Nation’s historical tradition of firearm regulation. Only then may a court conclude that the individual’s conduct falls outside the Second Amendment’s “unqualified command.”110Id. at 2129–30 (quoting Konigsberg v. State Bar of Cal., 366 U.S. 36, 49 n.10 (1961)). The Court’s quotation of Konigsberg here is misleading. The Court in Konigsberg compared the Second Amendment’s “unqualified command” with the restrictive reading of the right to bear arms in United States v. Miller, 307 U.S. 174 (1938) as an analogy for how the right to free speech is similarly not absolute, despite the First Amendment’s “unqualified terms.” Konigsberg, 366 U.S. at 49 n.10 (1961). The Court in Bruen, however, uses this quote as a semantic justification for a more expansive reading. This test dashed hopes that the Court would adopt a reasonability standard that states have largely applied to their own Second Amendment analogues. See Blocher, supra note 49, at 381–83; see also Winkler, supra note 50, at 687.

In applying this test, the Court stated that it would consider whether historical precedent from before, during, and relatively shortly after the founding demonstrates a “comparable tradition of regulation.”111Bruen, 142 S. Ct. at 2131–32 (citing Heller, 554 U.S. at 631). When comparing modern firearm laws and regulations to historical precedents, it is of course necessary to reason by analogy to determine whether the two are relatively similar.112Id. at 2132. Although the Bruen Court did not provide an exhaustive list of features that would render historical precedents relatively similar to modern laws, it provided two metrics: “how and why the regulations burden a law-abiding citizen’s right to armed self-defense.”113Bruen, 142 S. Ct. at 2133. Thus, for a historical precedent to support the constitutionality of a modern regulation, there must be (1) a comparable burden and (2) that burden must be comparably justified.114Importantly, the Court noted that, to successfully defend a regulation, the government must only find a proper “historical analogue, not a historical twin.” Id. at 2133 (emphasis in original); cf. Cass R. Sunstein, On Analogical Reasoning, 106 Harv. L. Rev. 741, 773 (1993) (“Everything is similar in infinite ways to everything else . . . . At the very least one needs a set of criteria to engage in analogical reasoning. Otherwise one has no idea what is analogous to what.”). For instance, there have long been prohibitions on carrying arms in sensitive places such as legislative assemblies, schools, and courthouses, so laws prohibiting carrying arms in such places, or even in newly defined sensitive places, are presumptively constitutional, so long as the sensitive place is analogous to those historically designated as such.115See id.; David B. Kopel & Joseph G.S. Greenlee, The “Sensitive Places” Doctrine: Locational Limits on the Right to Bear Arms, 13 Charleston L. Rev. 203, 227–36, 242–45 (2018); see also Heller, 554 U.S. at 626. New York’s licensing scheme, by contrast, could not be justified as analogous to these historical “sensitive places” laws because it generally banned citizens from carrying arms in any place “where people typically congregate,”116Bruen, 142 S. Ct. at 2133. meaning that entire cities would essentially be exempted from Second Amendment protection.117Id. at 2133–34. The Court also refused to allow the Second Amendment to be construed to apply only in the home. Id. at 2134 (“[T]he Second Amendment guarantees an ‘individual right to possess and carry weapons in case of confrontation,’ and confrontation can surely take place outside the home.”) (quoting Heller, 554 U.S. at 592).

Turing to New York’s proper-cause requirement, the Court stated that the plain text of the amendment protects ordinary citizens’ general right to carry handguns publicly for self-defense, emphasizing that confining the right to bear arms to the home would nullify half of the Second Amendment’s explicit protections—to not only “keep” but also “bear” arms.118Id. at 2134. The central right of the Second Amendment has been held to be the right to bear arms for self-defense in case of confrontation, which often will take place outside the home.119Id. at 2134–35; Heller, 554 U.S. at 592, 599. In assessing New York’s requirement that applicants for a license to carry a firearm in public show “proper cause”—as defined by the New York courts—the Court assessed a variety of sources that the respondents appealed to, dating from the 1200s to the early 1900s.120Bruen, 142 S. Ct. at 2135–36. The Court explained that, “when it comes to interpreting the Constitution, not all history is created equal.”121Id. at 2136. Therefore, even in light of the pre-existing right doctrine, historical evidence long-predating the enactment of the Second and Fourteenth Amendments will carry less weight than historical precedents closer in time to these enactments if legal conventions have changed in the intervening years.122Id. Thus, English practices traceable from the Middle Ages to the ratification of the Constitution will carry more weight than ancient practices that became obsolete before ratification.123Id. (citing Dimick v. Schiedt, 293 U.S. 474, 477 (1935)). Likewise, post-enactment history can be elucidating when “a regular course of practice” can settle the meaning of disputed terms and phrases.124Id. (quoting Chiafalo v. Washington, 140 S. Ct. 2316, 2326 (2020)); see also NLRB v. Noel Canning, 573 U.S. 513, 572 (2014) (Scalia, J., concurring) (“[W]here a governmental practice has been open, widespread, and unchallenged since the early days of the Republic, the practice should guide our interpretation of an ambiguous constitutional provision.”); The Federalist No. 37, at 179 (James Madison) (Buccaneer Books 1992) (“All new laws, though penned with the greatest technical skill, and passed on the fullest and most mature deliberation, are considered as more or less obscure and equivocal, until their meaning be liquidated and ascertained by a series of particular discussions and adjudications.”). However, when post-enactment precedents take effect long after ratification, they will be accorded less weight.125Bruen, 142 S. Ct. at 2137; cf. Sprint Commc’ns Co. v. APCC Servs., Inc., 554 U.S. 269, 312 (2008) (Roberts, C. J., dissenting) (“The belated innovations of the mid- to late-19th-century courts come too late to provide insight into the meaning of [the Constitution in 1787].”). With the parameters of its historical inquiry set, the Court proceeded to determine that the historical record the respondents compiled failed to demonstrate a historical analogue for New York’s firearm licensing scheme.126Bruen, 142 S. Ct. at 2138. That is, there was no historical tradition of limiting the public carry of firearms to citizens who could demonstrate a special need for self-defense, nor was there a historical tradition of broadly prohibiting the carry of commonly used firearms for self-defense.127Id.

A few key takeaways from the Court’s evaluation of this compendium of historical precedents will inform how a model gun control statute can be structured. First, the manner of public carry was historically subject to reasonable regulation—individuals could be restricted from carrying deadly weapons in a way that would be likely to terrorize others.128Id. at 2150. Second, states with surety laws129Surety statutes generally required certain individuals to post bond before carrying weapons in public. These were not the general bans absent a specific showing of a particular need as the New York statute was. Rather, these statutes targeted those threatening to do harm. Id. at 2148; see also Wrenn v. District of Columbia, 864 F.3d 650, 661 (D.C. Cir. 2017) (“[S]urety laws did not deny a responsible person carrying rights unless he showed a special need for self-defense. They only burdened someone reasonably accused of posing a threat. And even he could go on carrying without criminal penalty. He simply had to post money that would be forfeited if he breached the peace or injured others—a requirement from which he was exempt if he needed self-defense.”). provided financial incentives for responsible arms carrying, rather than directly restricting public carry.130Bruen, 142 S. Ct. at 2150. Third, states have historically been able to restrict or eliminate one kind of public carry—usually concealed carry—so long as they allowed the other type of carry—usually open carry.131Id. Fourth, the more widely enacted a type of statute is, the more likely the court is to uphold it. Thus, the relatively few historical examples prohibiting the carry of pistols—and in some cases all firearms—in towns, cities, and villages could not “overcome the overwhelming evidence of an otherwise enduring American tradition permitting public carry.”132Id. at 2154. Many of the statutes that prohibited or severely restricted the public carry of arms were enacted in the Western Territories prior to statehood. Id. The Court recognized two main defects in analogizing these statutes to modern legislation. First, the territorial populations that lived under these statutes was miniscule—less than one percent of the population at the time, showing that they were not widely adopted. Id. Second, the American territorial system was transitional and temporary, allowing for more improvisational territorial legislation that was short-lived and rarely subject to judicial scrutiny. Id. at 2155. Finally, as Kavanaugh’s concurrence underscores, the Court’s opinion does not jeopardize the existing “shall-issue” licensing regimes employed in forty-three states, only the “may-issue” regimes employed by six states and the District of Columbia.133Id. at 2161 (Kavanaugh, J., concurring). The states with “shall-issue” regimes are New York, California, Hawaii, Maryland, Massachusetts, and New Jersey. Id. at 2124. See also Thomson Reuters, 50 State Statutory Surveys: Right to Carry a Concealed Weapon, 0030 Surveys 32 (2022). The District of Columbia’s analogue to the “proper cause” standard at issue in Bruen has been enjoined since 2017. Wrenn, 864 F.3d at 668. The difference between these two is that when an applicant meets the statutory criteria in a shall-issue regime, they must be issued a license. Under a “may-issue” regime, however, even if an applicant meets the statutory criteria, a licensing officer has the discretion to refuse to issue a license, based on the difficult to meet “special need” requirement.134Bruen, 142 S. Ct. at 2123–21. Although the Court did not explicitly say that “shall-issue” regimes and “proper cause” requirements for licenses to carry firearms for self-defense are per se unconstitutional, it is difficult to see how either of these could be upheld.135See id. at 2138 n.9.

B.  Summary of Second Amendment Precedent

Before moving on to the model statute, a brief summary of the major limitations imposed by the foregoing trilogy of modern Second Amendment jurisprudence will prove helpful. First, the core right protected by the Second Amendment is an individual right to keep and bear arms for self-defense. Second, this right is effective against both the state and federal governments. Third, if the Second Amendment’s plain text—as interpreted by the Supreme Court—covers an individual’s conduct, it is presumptively protected, and the government must demonstrate that the law in question is analogous—though not necessarily identical—to a historical practice of firearms regulation. Fourth, when seeking a historical analogue to justify a modern regulation, not all history is created equal. Examples of post-ratification regulation that settle disputed terms and are relatively close in time to the adoption of the Bill of Rights can be particularly informative, as can evidence of English and Colonial practices that stayed in effect at least until ratification. Fifth, the more widespread a particular firearm regulation is, the more likely it is constitutional. Sixth, a legislature might well be able to ban or severely restrict either concealed carry or open, so long as they allow one of the two methods to remain legal. Finally, some types of firearm regulations are presumptively lawful—prohibitions on possession by felons and the mentally ill, laws against brandishing a firearm—while some are presumptively unlawful—shall-issue regimes, proper cause requirements.

As onerous as these requirements might appear to be, there is still a way for legislatures to assert meaningful control over the exercise of the Second Amendment, albeit with less free reign than they had previously been allowed. A systemic approach to gun ownership composed of rules that have historical analogues in the American legal tradition can be modeled on South Africa’s firearm licensing system. South Africa’s Firearms Control Act could provide a method to limit possession of high-capacity magazines while still allowing them to be owned for self-defense.

II.  SOUTH AFRICA’S GUN CONTROL SYSTEM

South Africa is fairly unique in its approach to firearms ownership in that a central focus of its firearm licensing system is to allow people the means to defend themselves.136Firearms Control Act 60 of 2000 pmbl. JSRSA (S. Afr.) (updated through 2014). Its licensing system is nevertheless comprehensive in spelling out the requirements for owning different categories of firearms and is fairly stringent in its requirements for firearm ownership in the first place—at least when compared with current law in the United States. Because the South African Model allows for a right to own a firearm for self-defense,137See id. at ch. 6 § 13. yet provides a comprehensive licensing scheme, it provides an ideal starting point for drafting a model statute for the United States.

The main feature of South Africa’s Firearms Control Act of 2000138The Act is designed around creating a comprehensive licensing system that requires a competency as well as a license for each firearm that a person owns. See id. at ch. 4 § 6(2) (“[N]o licence may be issued to a person who is not in possession of the relevant competency certificate.”).—which states could benefit from replicating—is a licensing system requiring citizens who wish to own a firearm to first obtain a competency certificate139Id. and then obtain a license specific to each firearm that they wish to own.140Id. at ch. 6 § 11 (“The Registrar must issue a separate licence in respect of each firearm licensed in terms of this Chapter.”). The type of firearm a citizen can own will depend on the type of license that they receive, which, in turn, depends on their purpose for owning the firearm. For instance, a citizen cannot obtain a semi-automatic rifle for occasional hunting or sports shooting because such a weapon is not necessary for that purpose.141See id. at ch. 15. Of course, a semi-automatic rifle could be used for occasional hunting or sports shooting, but the South African legislature presumably found that the potential danger of allowing more citizens to own semi-automatic rifles outweighed its utility for occasional hunting and sports shooting. This an important feature that could lawfully be replicated in the United States142See infra Section IV.B.2. to strike a balance between the states’ interest in public safety and the private interest in self-defense. Take high-capacity magazines, for instance. Some states have tried to outright ban them,143See, e.g., Safety for All Act, 2016 Cal. Legis. Serv. Prop. 63, § 6.1 (West), adding Cal. Penal Code § 32310(c)–(d) (Supp. 2020). but it is not clear that this is constitutional under Heller, McDonald, and Bruen.144Compare Duncan v. Becerra, 366 F. Supp. 3d 1131, 1143 (S.D. Cal. 2019) (“California’s § 32310 directly infringes Second Amendment rights . . . by broadly prohibiting common firearms and their common magazines holding more than 10 rounds, because they are not unusual and are commonly used by responsible, law-abiding citizens for lawful purposes such as self-defense.”), rev’d sub nom. Duncan v. Bonta, 19 F.4th 1087 (9th Cir. 2021), vacated, 142 S. Ct. 2895 (2022) (mem.), with Wiese v. Becerra, 306 F. Supp. 3d 1190, 1195 n.3 (E.D. Cal. 2018) (finding that California’s high-capacity magazine ban does not violate the Second Amendment), and Ass’n of N.J. Rifle & Pistol Clubs, Inc. v. Att’y Gen. of N.J., 910 F.3d 106, 118 (3d Cir. 2018) (finding that a New Jersey law banning high-capacity magazines “does not severely burden, and in fact respects, the core of the Second Amendment right”), abrogated by N.Y. Rifle & Pistol Ass’n v. Bruen, 142 S. Ct. 2111 (2022). South Africa’s Firearms Control Act could provide a method to limit possession of high-capacity magazines while still allowing them to be owned for self-defense uses.

Although South Africa’s system provides a good starting point for a model act, some areas will need modification to comport with U.S. constitutional standards. The main modifications are in the types of firearms that can be owned and the permit issuance requirements. Heller instructs that firearms in common use receive Second Amendment protection145See supra Section I.A.3. and Bruen indicates that “may issue” regimes are very likely per se unconstitutional.146See N.Y. State Rifle & Pistol Ass’n v. Bruen, 142 S. Ct. 2111, 2138 (2022). The primary modifications this Note proposes for its Model Act appear in Sections 2(b), 3, 5, and 6 in Part III below.

III.  THE MODEL FIREARMS CONTROL ACT

The following is the full text of the Model Firearms Control Act that this Note proposes the states adopt. This act is intended to supplement existing state firearms regulations by creating an individual licensing requirement.

A.  Chapter 1: Definitions, License Requirement, Eligibility Certificate

  • § 1 Definitions
    • (a) Designated Firearms Officer. A “Designated Firearms Officer” means a law enforcement official designated as such by state law.
    • (b) Accredited Hunting Association. An “Accredited Hunting Association” means an association meeting the criteria designated by state law.
    • (c) Accredited Sports Shooting Association. An “Accredited Sports Shooting Association” means an association meeting the criteria designated by state law.
  • § 2 License Requirement
    • (a) No person may possess a firearm without holding a license issued by the state for that firearm. A separate license is required for each firearm.
    • (b) No person may receive a license to possess a firearm without first having been issued an eligibility certificate.
    • (c) A Designated Firearms Officer may not issue an applicant a license to possess a firearm that is not legal to possess in the state within which the applicant resides.
    • (d) Familial transfer. Ownership of a firearm cannot be transferred from one person to another unless the transferee has a license to possess that firearm, even if the transferor and transferee are family members.
    • (e) Issuance. Upon meeting the criteria for any firearms license, the Designated Firearms Officer to whom the application has been delivered must issue the applicant the appropriate firearms license. Neither the Designated Firearms Officer, nor any other state or federal government employee or agent may prevent an applicant from delivering a valid application to the Designated Firearms Officer.
  • § 3 Eligibility Certificate
    • (a) Requirements. To receive an eligibility certificate, the applicant must:
      • (1) Complete an application delivered to a Designated Firearms Officer responsible for the area in which applicant resides;
      • (2) Provide a full set of the applicant’s fingerprints;
      • (3) Be eighteen years old or older;
      • (4) Be lawfully present in the United States;
      • (5) Not suffer from a mental illness that renders the applicant a danger to himself or others;
      • (6) Never have been convicted of a crime punishable by a year or more of imprisonment;
      • (7) Never have been convicted of a crime involving:
        • (A) Fraud in relation to—or supplying false information for the purpose of—obtaining an eligibility certificate, license, permit, or authorization to possess a firearm in terms of this Act or a previous law; or
        • (B) Domestic violence.
      • (8) Not be addicted to any substance that has an intoxicating effect; and
      • (9) Pass a firearms safety course as prescribed by state law.
    • (b) Issuance. Upon the applicant’s completion of the above requirements, the Designated Firearms Officer to whom an application has been delivered must issue a qualified applicant an eligibility certificate within thirty days of delivery.
    • (c) Denial pending investigation. If the Designated Firearms Officer has a well-founded doubt that an applicant has not met one or more of the eligibility requirements, the Designated Firearms Officer can deny an applicant an eligibility certificate for up to thirty days, during which time he or she may conduct a further investigation to determine whether the applicant has met the requirements to receive an eligibility certificate. After thirty days, the Designated Firearms Officer must either:
      • (1) Issue the eligibility certificate if the applicant meets the necessary criteria; or
      • (2) Provide the applicant with the reason for denying the certificate.

      If the Designated Firearms Officer has a well-founded doubt as to the mental stability of an applicant, the Designated Firearms Officer has the discretion to require an applicant to undergo a screening by a licensed psychologist or licensed psychiatrist before issuing an eligibility certificate contingent on the psychologist or psychiatrist’s determination that the applicant is of stable mental condition.

B.  Chapter 2: Types of Licenses and Use of Firearms

  • § 4 License to Possess a Firearm for Self-Defense
    • (a) Firearms that may be possessed for self-defense. A person can receive a license to possess the following firearms for self-defense:
      • (1) A handgun that is not fully automatic; or
      • (2) A rifle or shotgun that is not semi-automatic or fully automatic.
    • (b) Issuance. A license to possess such a firearm for self-defense must be issued to any natural person who
      • (1) Holds a valid eligibility certificate; and
      • (2) Applies for a license to possess a firearm for self-defense.
    • (c) Limits. No person may possess more than two licenses under this section.
  • § 5 License to Possess a Restricted Firearm for Self-Defense
    • (a) Restricted firearms defined. The following are considered “restricted firearms” for the purpose of this section:
      • (1) A rifle or shotgun that accepts detachable magazines and is semi-automatic but not fully automatic.
    • (b) Requirements to issue a license to possess a restricted firearm for self-defense. A license to possess a firearm for self-defense must be issued to any natural person who
      • (1) Holds a valid eligibility certificate;
      • (2) Applies for a license to possess a restricted firearm for self-defense; and
      • (3) Shows cause why the particular restricted firearm for which a license is sought serves a self-defense need that a non-restricted firearm cannot serve.
    • (c) Basis for denial. The Designated Firearms Officer reviewing the application to possess a restricted firearm for self-defense must provide an objectively reasonable basis, based on clear and convincing evidence, to deny an application for lack of cause under section 5(b)(3).
  • § 6 License to Carry a Concealed Firearm for Self-Defense
    • (a) Concealed carry defined. “Concealed carry” means carrying a firearm on the person of the license holder in a way that is not visible to others.
    • (b) Requirements to issue a license to carry a concealed firearm for self-defense. A license to possess a firearm for self-defense must be issued to any natural person who
      • (1) Is twenty-one years old or older;
      • (2) Holds a valid eligibility certificate; and
      • (3) Completes an application delivered to a Designated Firearms Officer responsible for the area in which the applicant resides.
    • (c) Arms that may be possessed for concealed carry. A person who holds a license to carry a concealed firearm can carry any handgun that is concealable on the person, is not fully automatic, and that the person has a license to possess.
  • § 7 License to Openly Carry a Firearm for Self-Defense147Either this section or section 5 can be deleted at the determination of the state legislature, but one type of public carry—either open or concealed—must be permitted. See Bruen, 142 S. Ct. at 2150.
    • (a) Openly carry defined. “Openly carry” means carrying a firearm on the person of the license holder that is not concealed.
    • (b) Requirements to issue a license to openly carry a firearm for self-defense. A license to openly carry a firearm for self-defense must be issued to any natural person who
      • (1) Is twenty-one years old or older;
      • (2) Holds a valid eligibility certificate; and
      • (3) Completes an application delivered to a Designated Firearms Officer responsible for the area in which applicant resides.
    • (c) Arms that may be possessed for open carry. A person who holds a license to openly carry a firearm can openly carry any handgun that is not fully automatic and that the person has a license to possess.
  • § 8 License to Possess a Firearm for Occasional Hunting or Occasional Sports Shooting
    • (a) Purpose. The purpose of this section is to provide responsible, law-abiding citizens access to ordinary firearms for such lawful activities as hunting, target practice, and sports shooting.148The terms “occasional hunting” and “occasional sports shooting” remain undefined because definition is not necessary. Section 8 is rather permissive in providing access to ordinary firearms (as opposed to dangerous and unusual firearms) to any person who can obtain an eligibility certificate.
    • (b) Persons eligible under this section. Any person who holds a valid eligibility certificate can receive a license to possess a firearm for occasional hunting or sports shooting.
    • (c) Arms that may be possessed for occasional hunting or sports shooting. A person may receive a license to possess the following firearms for occasional hunting or occasional sports shooting:
      • (1) A rifle or shotgun that is not semi-automatic or fully automatic; and
      • (2) A handgun that is not fully automatic.
  • § 9 License to Possess a Firearm for Dedicated Hunting or Dedicated Sports Shooting
    • (a) Dedicated hunter defined. A “dedicated hunter” means a person who regularly participates in hunting activities and who is a member of an Accredited Hunting Association.
    • (b) Dedicated sports shooter defined. A “dedicated sports shooter” means any person who regularly participates in sports shooting and who is a member of an Accredited Sports Shooting Association.
    • (c) A person who is a dedicated hunter or a dedicated sports shooter can receive a license to possess the following firearms for dedicated hunting or dedicated sports shooting:
      • (1) A rifle or shotgun that is not fully automatic; and
      • (2) A handgun that is not fully automatic.

    C.  Chapter 3: Use and Transportation of Firearms

    • § 10 Use of Firearms. A person may use a firearm for which the person holds a valid license where it is safe and lawful to do so.
    • § 11 Transportation of Firearms. A person lawfully possessing a firearm can transport that firearm in any manner that is consistent with state law.

    IV.  CONSTITUTIONALITY

    The Model Act that this Note proposes is designed to survive judicial review by United States courts, rather than to be considered constitutional in an abstract sense. The object of the application of Second Amendment jurisprudence here is “the prediction of the incidence of the public force through the instrumentality of the courts.”149Justice O. W. Holmes, Address at the Boston University School of Law: The Path of the Law 457 (Jan. 7, 1897), in 10 Harv. L. Rev. 457 (1897). As such, this Section argues that the confluence of Second Amendment doctrine and practical considerations will allow the Model Act to remain “lawful” in the realist sense.150See id. at 461 (“The prophecies of what the courts will do in fact, and nothing more pretentious, are what I mean by the law.”).

    A.  “Longstanding Prohibitions”

    The Heller Court enumerated in dictum certain restrictions on the right to bear arms that its common use doctrine did not place in jeopardy.

    [N]othing in our opinion should be taken to cast doubt on longstanding prohibitions on the possession of firearms by felons and the mentally ill, or laws forbidding the carrying of firearms in sensitive places such as schools and government buildings, or laws imposing conditions and qualifications on the commercial sale of arms.151District of Columbia v. Heller, 554 U.S. 570, 626–27 (2008). Unfortunately, the Heller Court provided no historical basis for these restrictions, so the Heller opinion itself is of no use in finding a historical precedent for these “longstanding prohibitions.” Id. at 626–27; see Waldman, supra note 7, at 126 (“This eminently sensible list barges into the text, seemingly from nowhere.”). In his McDonald dissent, Justice Breyer succinctly summarizes the odd nature of this list of “acceptable” regulations.
    [T]he Court has haphazardly created a few simple rules, such as that it will not touch “prohibitions on the possession of firearms by felons and the mentally ill,” “laws forbidding the carrying of firearms in sensitive places such as schools and government buildings,” or “laws imposing conditions and qualifications on the commercial sale of arms.” But why these rules and not others? Does the Court know that these regulations are justified by some special gun-related risk of death? In fact, the Court does not know. It has simply invented rules that sound sensible without being able to explain why or how Chicago’s handgun ban is different.
    McDonald v. City of Chicago, 561 U.S. 742, 925 (2010) (Breyer, J., dissenting) (citations omitted) (quoting Heller, 554 U.S. at 626–27).

    Although somewhat reassuring at the time the Heller decision was handed down, Bruen and McDonald have not given this assertion clear majority support. First, in McDonald, only Chief Justice Roberts and Justices Scalia and Kennedy joined Justice Alito’s endorsement of this list of presumptively lawful restrictions.152McDonald, 561 U.S. at 786 (plurality opinion). Next, in Bruen, this passage was omitted entirely from the majority opinion, appearing only in Justice Kavanaugh’s concurrence.153Bruen, 142 S. Ct. at 2162 (Kavanaugh, J., concurring). Perhaps, then, this ipse dixit of “longstanding prohibitions” will not carry any special favor with the Court in the future and sections 3(a)(1–7) of the Model Act will have to be justified under Bruen’s historical test.

    1.  Prohibition on Firearm Possession by Felons

    When applying Bruen’s historical test to section 3(a)(7) of the Model Act—which denies eligibility certificates to felons—the first question is whether the Second Amendment’s plain text covers an individual’s conduct.154Id. at 2126. To conclude that the Second Amendment does not cover this conduct requires reliance more on dicta from Heller, McDonald, and Bruen, as well as the majority’s focus on the rights of law-abiding citizens in these cases,155Of course, the law-abiding nature of the litigants in Heller, McDonald, and Bruen was never in question, limiting the persuasiveness of this argument. rather than a formally applied Bruen analysis. In United States v. Minter, for instance, a district court was faced with a challenge to the constitutionality of a federal law that makes possession of firearms and ammunition by convicted felons illegal.156United States v. Minter, 635 F. Supp. 3d 352, 354 (M.D. Pa. 2022). The district court reasoned that “the Supreme Court in Bruen ha[d] already signaled the answer to this question,” and concluded that “the Bruen Court’s decision did not undermine Heller’s statement,” emphasizing that the Second Amendment protects the “right of law-abiding, responsible citizens to use arms for self-defense.”157Id. at 358 (quoting Bruen, 142 S. Ct. at 2131) (emphasis in original). Several other district courts have considered the constitutionality of a felon-in-possession laws post-Bruen, many of which have concluded that the Second Amendment’s plain text does not cover this conduct.158See, e.g., United States v. Ingram, 623 F. Supp. 3d 660, 664 (D.S.C. 2022) (“[S]imilar discussion regarding felon-in-possession and comparable statutes appears in three different opinions: Heller, McDonald, and Bruen. By distinguishing non-law-abiding citizens from law-abiding ones, the dicta in Heller and McDonald clarifies the bounds of the plain text of the Second Amendment.”); United States v. Jackson, No. 21-51, 2022 U.S. Dist. LEXIS 164604, at *3 (D. Minn. Sept. 13, 2022) (“In Bruen, the Court again stressed that Heller and McDonald remain good law. Justice Kavanaugh, joined by Chief Justice Roberts, stated that Bruen does not disturb what the Court has said in Heller about the restrictions imposed on possessing firearms . . . .”); United States v. Hill, 629 F. Supp. 3d 1027, 1029–30 (S.D. Cal. 2022); United States v. Siddoway, No. 21-cr-00205, 2022 U.S. Dist. LEXIS 178168, at *3–5 (D. Idaho Sept. 27, 2022); United States v. Burrell, No. 21-20395, 2022 U.S. Dist. LEXIS 161336, at *6–7 (E.D. Mich. Sept. 7, 2022). However, this reliance on dicta might not be enough to avoid the historical inquiry that Bruen demands.159See, e.g., United States v. Coombes, 629 F. Supp. 3d 1149, 1154–56 (N.D. Okla. 2022) (concluding that the Second Amendment does not categorically exclude convicted felons from “the people”).

    If the Second Amendment’s plain text is interpreted to include convicted felons in its reference to “the people,”160Id. Bruen’s historical test would require the government to determine whether section 3(a)(7) is “consistent with this Nation’s historical tradition of firearm regulation.”161Bruen, 142 S. Ct. at 2126. Because the earliest felon-disarmament laws date from the twilight of the nineteenth century and the early part of the twentieth century,162Carlton F.W. Larson, Four Exceptions in Search of a Theory: District of Columbia v. Heller and Judicial Ipse Dixit, 60 Hastings L.J. 1371, 1376 (2009). an earlier historical analogue must be identified. One possible analogue is some American Colonies’ practice of attainder, which was utilized to disarm “disaffected” or “delinquent” individuals.163See Thomas R. McCoy, The Collateral Consequences of a Criminal Conviction, 23 Vand. L. Rev. 929, 942–49, 1080–82 (1970); 1 Journals of the Provincial Congress, Provincial Convention, Committee of Safety and Council of Safety of the State of New York 149–50 (Albany, Thurlow Weed 1842); see also Stephen P. Halbrook, The Founders’ Second Amendment: Origins of the Right to Bear Arms 117–18 (2008). Although a bill of attainder would surely constitute a due process violation today, the colonial practice of attainder is still sufficiently analogical to felon-in-possession laws because it is an example of state legislatures disarming individuals considered dangerous.164See Coombes, 629 F. Supp. 3d at 1157–58. Additionally, a colonial New York law prohibited convicted felons from owning property or chattels, implicitly prohibiting them from owning firearms.165See Howard Itzkowitz & Lauren Oldak, Note, Restoring the Ex-Offender’s Right to Vote: Background and Developments, 11 Am. Crim. L. Rev. 721, 725 n.33 (1973); 1 The Colonial Laws of New York: From the Year 1664 to the Revolution 145 (Albany, James B. Lyon 1894); Coombes, 629 F. Supp. 3d at 1158–59. Finally, a few historical examples of proposals made during the ratification process show that the founders did not intend to confer the right to bear arms on convicted felons.166See Coombes, 629 F. Supp. 3d at 1158–59. Proposals from Anti-Federalists in Pennsylvania,1672 Bernard Schwartz, The Bill of Rights: A Documentary History 665 (1971). Samuel Adams in Massachusetts,168Heller, 554 U.S. at 716 (Breyer, J., dissenting). and delegates from New Hampshire1691 Jonathan Elliot, The Debates in the Several State Conventions on the Adoption of the Federal Constitution 326 (Philadelphia, J. B. Lippincott Co. 2d ed. 1836) (The proposed amendment provided that “Congress shall never disarm any citizen, unless such as are or have been in actual rebellion.”). all show that the framers thought of the Second Amendment as recognizing the right of law-abiding citizens to bear arms. One of these proposals, for instance, provided that “no law shall be passed for disarming the people or any of them unless for crimes committed, or real danger of public injury from individuals.”170Schwartz, supra note 167. Although these were only proposals, they are still helpful because they illustrate how Americans at the time understood the government’s authority to limit firearm possession. These proposals’ rejection does not necessarily show that early Americans objected to such limitations on the right to bear arms and could simply be a result of a lack of Federalist support.

    Although the historical precedents identified here are not overly persuasive, they have thus far been sufficient for many federal courts that have heard challenges to the federal felon-in-possession law and entertained the question of whether it is consistent with this Nation’s historical tradition of firearm regulation.171See, e.g., Coombes, 629 F. Supp. 3d at 1158–59; United States v. Collette, 630 F. Supp. 3d 841, 846 (W.D. Tex. 2022); United States v. Charles, 633 F. Supp. 3d 874, 878–79 (W.D. Tex. 2022); United States v. Price, 635 F. Supp. 3d 455, 458 (S.D.W. Va. 2022); United States v. Cockerham, No. 21-cr-6, 2022 U.S. Dist. LEXIS 164702, at *3–4 (S.D. Miss. Sept. 13, 2022). Taking a realist view, this could simply be because the judiciary is staffed by those “who know too much to sacrifice good sense to a syllogism”172O. W. Holmes, Jr., The Common Law 36 (Boston, Little, Brown, & Co. 1881). and are unwilling to invalidate a law that is so sensible on its face. Even the Supreme Court, staffed as it is today, might not invalidate such laws. Assuming Justice Kavanaugh’s concurring opinion in Bruen to be an honest representation of how he (and Chief Justice Roberts, who joined his concurrence) will vote in the future, the Heller Court’s enumeration of presumptively lawful regulations will not be stripped of all persuasive effect.173N.Y. Rifle & Pistol Ass’n v. Bruen, 142 S. Ct. 2111, 2161 (2022) (Kavanaugh, J., concurring). After all, it is hard to imagine that Justices Sotomayor, Kagan, or Jackson would not endorse the “longstanding prohibitions” passage from Heller. Thus, we can expect laws that prohibit felons from possessing firearms—and section 3(a)(6) of the Model Act—will not be invalidated by the Supreme Court, even if on practical rather than doctrinal considerations.

    2.  Prohibiting Firearm Possession for Certain Non-Felonies

    Possibly more challenging, however, is section 3(a)(7), which denies eligibility certificates to individuals convicted of fraud for the purpose of obtaining a firearm; unlawful use or handling of a firearm; or domestic violence. After deciding Bruen, the Supreme Court vacated and remanded for further consideration a circuit court decision rejecting a challenge to a Massachusetts law similar to section 3(a)(7) of the Model Act.174Morin v. Lyver, 143 S. Ct. 69, 69 (2022). The law in question prohibited the plaintiff from receiving a license to carry a firearm in public because he had been convicted of attempting to carry a pistol without a license and of possession of an unregistered firearm in the District of Columbia.175Morin v. Lyver, 13 F.4th 101, 102–03 (1st Cir. 2021), vacated, 143 S. Ct. 69 (2022). Although these convictions were misdemeanors, Massachusetts law denied public carry licenses to “persons who had, ‘in any state or federal jurisdiction, been convicted’ of ‘a violation of any law regulating the use, possession, ownership, transfer, purchase, sale, lease, rental, receipt or transportation of weapons or ammunition for which a term of imprisonment may be imposed.’ ”176Id. at 103 (quoting Mass. Gen. Laws ch. 140, § 131(d)(i)(D) (2008)). In upholding the Massachusetts law, the circuit court applied intermediate scrutiny, which the Supreme Court has since rejected as inappropriate for Second Amendment analysis.177Bruen, 142 S. Ct. 2111 at 2129–30. However, in a similar post-Bruen case, the Fifth Circuit upheld a law prohibiting possession of a firearm by persons under indictment,178United States v. Avila, No. 22–50088, 2022 U.S. App. LEXIS 35321, at *1 (5th Cir. Dec. 21, 2022); see also United States v. Rowson, No. 22 Cr. 310, 2023 U.S. Dist. LEXIS 13832, at *1 (S.D.N.Y. Jan. 26, 2023). and several district courts have upheld laws prohibiting possession of firearms by felons,179See, e.g., United States v. Kays, 624 F. Supp. 3d 1262, 1265 (W.D. Okla. 2022); United States v. Price, 635 F. Supp. 3d 455, 466–67 (S.D.W. Va. 2022); United States v. Minter, 635 F. Supp. 3d 352, 354 (M.D. Pa. 2022); District of Columbia v. Heller, 554 U.S. 570, 716 (2008) (Breyer, J., dissenting). domestic violence misdemeanants,180United States v. Nutter, 624 F. Supp. 3d 636, 644–45 (S.D.W. Va. 2022). Infamously, however, the Fifth Circuit recently held that the federal law prohibiting possession of firearms by persons under a domestic violence restraining order is unconstitutional because it does not fit “within our Nation’s historical tradition of firearm regulation.” United States v. Rahimi, 61 F.4th 443, 460 (5th Cir. 2023), cert granted, 143 S. Ct. 2688 (2023). and unlawful users of controlled substances.181United States v. Daniels, 610 F. Supp. 3d 892, 897 (S.D. Miss. 2022), rev’d, 77 F.4th 337 (5th Cir. 2023). Although some district courts have held similar laws unconstitutional,182United States v. Hicks, No. W:21-CR-00060, 2023 U.S. Dist. LEXIS 35485, at *17–18  (W.D. Tex. Jan. 9, 2023) (holding a law prohibiting firearm possession while under a felony indictment unconstitutional); United States v. Quiroz, 629 F. Supp. 3d 511, 511–12 (W.D. Tex. 2022); Price, 635 F. Supp. 3d at 457 (holding a law prohibiting possession of a firearm with an altered, obliterated, or removed serial number unconstitutional); United States v. Perez-Gallan, 640 F. Supp. 3d 697,  716 (W.D. Tex. 2022) (holding a federal statute prohibiting firearm possession by those subject to restraining order related to domestic violence unconstitutional). there is, as of yet, no circuit precedent invalidating these laws.

    In addition to the weight of circuit precedent, the plain text of Heller supports the conclusion that “prohibitions on the possession of firearms by felons and the mentally ill”183Heller, 554 U.S. at 626. and similar measures are “presumptively lawful.”184Id. at 627 n.26. However, if the Court determines that Bruen abrogates the “longstanding prohibitions” dictum from Heller, a historical analogue will have to be found.185N.Y. Rifle & Pistol Ass’n v. Bruen, 142 S. Ct. 2111, 2133 (2022). Bruen provides two metrics to be considered in determining whether a regulation is relevantly similar to a historical analogue: (1) how they burden a “law-abiding citizen’s right to armed self-defense,” and (2) why they burden that right.186Id. at 2132–33. These are not the only metrics that could render a historical analogue “relatively similar,” but they are the only metrics the Court explicitly mentioned.

    Under these metrics, section 3(a)(7) of the Model Act could escape invalidation on the same basis as felon-in-possession laws187See supra Section IV.A.1.: because it burdens a law-abiding citizen’s right of lawful self-defense in a way similar to, and for a reason practically identical to, the colonial practice of disarming “disaffected” or “delinquent individuals” through attainder,188See McCoy, supra note 163, at 942–49. and the colonial practice of prohibiting convicted felons from owning chattels, including firearms.189Itzkowitz et al., supra note 165, at 721, 725 n.33.

    First, the burden is similar because a regulation prohibiting possession of firearms to certain classes of misdemeanants does not actually burden the right any more than a law prohibiting a felon’s possession of firearms. The right described in Bruen is one belonging to law-abiding citizens, not citizens convicted of felonies.190Bruen, 142 S. Ct. at 2138 n.9. Individuals convicted of a felony or misdemeanor domestic violence; unlawful use or handling of a firearm; or fraud for the purpose of unlawfully obtaining a firearm are by definition not law-abiding.191This does not mean, however, that any violation of the law will result in a forfeiture of Second Amendment rights. Section 3(a)(7) contemplates specific violations of law that tend to show that allowing that person to own a firearm would be dangerous to themselves, to the public, or both. To the contrary, those individuals would be showing that they are willing to commit serious violent crimes—domestic violence—or crimes showing that they are not safe users of firearms. Second, the reason for the restrictions in section 3(a)(7) of the Model Act are identical to the reasons for the colonial practice of prohibiting dangerous individuals from owning firearms: to ensure those bearing arms are responsible, safe, and law-abiding. In discussing the regulations in shall-issue licensing regimes, Bruen acknowledges that regulations designed “to ensure only that those bearing arms in the jurisdiction are, in fact, ‘law-abiding, responsible citizens’ ” are constitutional.192Bruen, 142 S. Ct. at 2138 n.9. Because section 3(a)(7) is closely analogous to the presumptively lawful measures expounded in Heller,193District of Columbia v. Heller, 554 U.S. 570, 626–27 (2008); see also McDonald v. City of Chicago, 561 U.S. 742, 786 (2010); Bruen, 142 S. Ct. at 2162 (Kavanaugh, J., concurring). it is likely to be held constitutional.

    Another potentially problematic provision is section 3(a)(8), which does not allow individuals addicted to narcotics to obtain an eligibility certificate that is a prerequisite to possession of any firearm. In 2023, a federal court in Oklahoma ruled that the federal statute prohibiting possession of firearms by users of substances made illegal by the federal Controlled Substances Act19418 U.S.C. § 922(g)(3). was unconstitutional.195United States v. Harrison, No. CR-22-00328, 2023 U.S. Dist. LEXIS 18397, at *51 (W.D. Okla. Feb. 3, 2023) (concluding that the statute forbidding possession of firearms by unlawful drug users violates the Second Amendment). However, this ruling is far from sounding the death knell for laws prohibiting possession of firearms by drug addicts. Even if this position was adopted by the circuit courts or the Supreme Court, it would not invalidate section 3(a)(8) because that section only prohibits individuals who are addicted to, rather than mere users of, intoxicating substances from obtaining eligibility certificates. This is intended to prevent individuals whose mental stability would be regularly impaired by the abuse of drugs or alcohol from possessing firearms and would not apply to an occasional marijuana user. Section 3(a)(8) is therefore very similar to a law prohibiting possession by those with mental illnesses, as described in Heller as presumptively lawful.196Heller, 554 U.S. at 626. These presumptively lawful restrictions were also endorsed by two justices in the majority in Bruen and would likely also be endorsed by the three dissenting justices. See Bruen, 142 S. Ct. at 2162 (Kavanaugh, J., concurring).

    One final challenge section 3(a)(8) might face is that it is unconstitutional under Robinson v. California.197Robinson v. California, 370 U.S. 660 (1962). In Robinson, the Court held that a law criminalizing being addicted to the use of narcotics was cruel and unusual punishment under the Eight Amendment.198Id. at 666; U.S. Const. amend. VIII. This comparison is, however, inapposite. Section 3(a)(8) does not criminalize drug addiction; it only prevents drug addicts from arming themselves—for their own safety and the safety of the general public. It is fundamentally no different from making it illegal for blind persons to drive. Moreover, the Model Act does not prevent persons who were once addicted to drugs but are no longer addicted from obtaining an eligibility certificate. Thus, section 3(a)(8) falls far short of being a punishment at all, much less a cruel and unusual one.

    B.  Regulation of Different Classes of Arms

    1.  Purpose-Based Licensing

    The defining characteristic of the Model Act, in accordance with South Africa’s licensing system, is how it ties the ownership of a firearm to its use by only allowing ownership of firearms that are suited to the purpose for which the license is sought. Although access to certain firearms, such as semi-automatic rifles, is restricted under the Model Act, they are not entirely banned. This is done in an attempt to limit access to especially dangerous firearms while acknowledging the reality that a blanket ban on assault weapons might not be held constitutional by the current Supreme Court because of the inherent difficulty in finding a historical precedent regulating distinctly modern arms.199See Miller v. Bonta, No. 19-cv-01537, 2023 U.S. Dist. LEXIS 188421, at *97 (S.D. Cal. Oct. 19, 2023). Moreover, given the Supreme Court’s current 6-3 conservative supermajority, a blanket ban would create a risk of creating a dangerous precedent. If the Supreme Court invalidated an assault weapons ban, future Justices who might not consider such a ban unconstitutional per se might nevertheless feel duty-bound to apply Supreme Court precedent faithfully.

    The requirements for a license to possess a firearm for self-defense described in section 4(b) of the Model act would likely be found constitutional under Bruen because it is very closely analogous to the “shall-issue” carry licensing system in place in the vast majority of states.200Bruen, 142 S. Ct. at 2162. Bruen held only that the discretion afforded to licensing officials in the states with “may-issue” regimes is unconstitutional,201Id. at 2156. and did not jeopardize the licensing requirements that are employed in forty-three states.202Id. at 2138 n.9; see also id. at 2161 (Kavanaugh, J., concurring) (“[T]he Court’s decision does not affect the existing licensing regimes—known as ‘shall-issue’ regimes—that are employed in 43 States.”); id. at 2162 (Kavanaugh J. concurring) (“[T]he Second Amendment allows a ‘variety’ of gun regulations.”) (citing District of Columbia v. Heller, 554 U.S. 570, 636 (2008)). The Court explained that “nothing in our analysis should be interpreted to suggest the unconstitutionality of the 43 States’ ‘shall-issue’ licensing regimes, under which ‘a general desire for self-defense is sufficient to obtain a [permit].’ ”203Bruen, 142 S. Ct. at 2138 n.9 (quoting Drake v. Filko, 724 F.3d 426, 442 (3d Cir. 2013) (Hardiman, J., dissenting)). Moreover, the Court used the fact that “shall-issue” licensing regimes are in place in the vast majority of states to bolster its argument that New York’s “may-issue” regime was unconstitutional.204Id. at 2123–24. The Court further explained that states are free to “require applicants to undergo a background check or pass a firearms safety course,” and that these measures “are designed to ensure only that those bearing arms in the jurisdiction are, in fact, ‘law-abiding, responsible citizens.’ ”205Id. at 2138 n.9 (quoting Heller, 554 U.S. at 635). Section 4(b) of the Model Act follows the convention of a “shall-issue” licensing regime, but applies to firearm ownership for self-defense in general, not just to public carry. Regulations such as section 4(b) of the Model Act, just like regulations that the court mentioned,206Id. serve only to ensure that firearm owners are law-abiding, responsible citizens.207Id. Thus, section 4(b) burdens the right of law-abiding citizens to keep and bear arms for self-defense to a similar extent, and for the very same purpose, as the public carry licensing requirements in effect in forty-three states.208See id. at 2123–24; Thomson Reuters, 50 State Statutory Surveys: Right to Carry a Concealed Weapon, 0030 Surveys 32 (2022). Although the modern prevalence of the licensing requirement might not be doctrinally relevant, practically speaking, 4(b) would be unlikely to be invalidated unless a court were either willing to invalidate the widespread practice of public carry licensing or unwilling to accept licensing for firearm ownership in general.

    2.  High-Capacity Magazines and Semi-Automatic Rifles

    Perhaps the most difficult constitutional question in this area is whether states can ban specific types of arms. The Supreme Court has not given clear guidance on these issues in any of its decisions, resulting in discordant lower court rulings on the issues of high-capacity magazine bans209Compare Ass’n of N.J. Rifle & Pistol Clubs Inc. v. Att’y Gen. of N.J., 910 F.3d 106, 117 (3d Cir. 2018) (“The Act [banning high-capacity magazines] does not severely burden the core Second Amendment right to self-defense in the home . . . .”), abrogated by N.Y. Rifle & Pistol Ass’n v. Bruen, 142 S. Ct. 2111 (2022), with Duncan v. Becerra, 970 F.3d 1133, 1169 (9th Cir. 2020) (“California’s near-categorical ban of LCMs [Large Capacity Magazines] infringes on the fundamental right to self-defense.”), vacated sub nom. Duncan v. Bonta, 142 S. Ct. 2895 (2022), and Duncan v. Bonta, No. 17-CV-1017, 2023 U.S. Dist. LEXIS 169577, at *4 (S.D. Cal., Sept. 22, 2023) (“There is no American tradition of limiting ammunition capacity . . . .”), appeal docketed, No. 23-55805, 2023 U.S. App. LEXIS 25723 (9th Cir. Sept. 28, 2023). and assault weapon bans.210Compare Bianchi v. Frosh, 858 Fed. App’x 645, 646 (per curiam) (4th Cir. 2021) (affirming district court’s dismissal of challenge to Maryland’s assault weapons ban for failure to state a claim on which relief may be granted), vacated, 142 S. Ct. 2898 (2022), with Miller v. Bonta, 542 F. Supp. 3d 1009, 1021 (S.D. Cal. 2021) (“The overwhelming majority of citizens who own and keep the popular AR-15 rifle and its many variants do so for lawful purposes, including self-defense at home. Under Heller, that is all that is needed.”), vacated, No. 21-55608, 2022 U.S. App. LEXIS 21172 (9th Cir. Aug. 1, 2022). Heller states that the Second Amendment protects individual ownership of the types of firearms in common use, and that this protection means states cannot outright ban handgun ownership, but it is not clear how expansively “common use” (or for that matter, the “type” of a firearm) will be interpreted. Several states have attempted to prohibit possession of high-capacity magazines and assault weapons such as the AR-15;211See, e.g., Md. Code Ann., Pub Safety § 5-101 et seq. (West 2022). however, until the issue is squarely addressed by the Court, states will have to operate based on discordant lower federal court decisions.

    Adding to this opacity is the term “assault weapon” itself. “Assault weapon” is a legal term that is not uniformly defined by legislatures.212Compare Md. Code Ann., Pub. Safety § 5-101(r)(2) (West 2022) (defining assault weapons by enumerating specific makes and models), with Cal. Penal Code § 30515 (West 2012 & Supp. 2020) (defining assault weapons as semi-automatic rifles with detachable magazines and certain combinations of features), and Public Safety and Recreational Firearms Use Protection Act, H.R. 4296, 103d Cong. (1994) (repealed 2004) (defining an assault weapon as either one of a list of enumerated makes and models or as having a combination of specific features). Some states define these weapons by its semi-automatic213Semi-automatic means having a mechanism for self-loading, but not continuous firing. That is, semi-automatic firearms allow for one shot per trigger pull without requiring manual operation of the firearm between shots. functioning in combination with features like flash hiders,214Used for reducing the amount of muzzle flash produced by a firearm upon discharging. pistol grips,215A grip shaped like the butt of a pistol. and adjustable stocks.216See, e.g., Cal. Penal Code § 30515 (West 2012 & Supp. 2020). Regardless of how they are defined, however, the most functionally important aspects of assault weapons are that they are semi-automatic rifles and accept detachable magazines.217Detachable magazines can be removed from a firearm without disassembly, allowing for much faster reloads. Assault weapons typically fire an intermediate rifle cartridge—a cartridge that is less powerful than a full-power rifle cartridge but more powerful than a pistol cartridge—making for a light and easy-to-use weapon with low recoil.218Phil Klay, Uncertain Ground: Citizenship in an Age of Endless, Invisible War 152–53 (2022). Perhaps the most common cartridge used in assault weapons in the United States is the 5.56 x 45mm NATO round.219Id. Although the projectile weighs only one tenth of an ounce, it is capable of traveling at up to 3,200 feet per second—almost triple the speed of sound—making for a rather destructive weapon.220Id. at 152. These light but fast bullets have the distinct advantage of producing low recoil while inflicting more damage than would be expected from its muzzle energy alone. Id. at 153. A 5.56 mm bullet from an AR-15 will begin tumbling and fragmenting at approximately eleven centimeters into the body, causing hydrostatic shock that can sever muscle tissue and burst apart organs. Id. Because the functionally important aspect of an assault weapon is that it is a semi-automatic rifle that accepts detachable magazines, the Model Act addresses these features specifically rather than fussing over the minute details of weapons that make little functional difference.

    In acknowledgement of the constitutional invalidation risk that an outright ban on high-capacity magazines or assault weapons poses,221Part of the danger of this overruling risk is that the Court could have occasion to announce a sweeping decision. the Model Act takes an intermediate approach limiting, but not prohibiting, access to assault weapons and does not attempt to regulate magazine capacity. Under sections 4 and 5 of the Model Act, citizens cannot be granted a license to possess a semi-automatic rifle or shotgun—classified as restricted firearms under section 5(a)—for self-defense unless they show the particular restricted firearm for which they are seeking a license serves an important purpose for which a non-restricted firearm is insufficient. For instance, if a person lives on a property with large open fields, a handgun might not be sufficient for self-defense because it is difficult to shoot accurately over a long distance and a manually operated rifle222A manually operated rifle is one that requires manual manipulation of the rifle’s action to chamber a new round and fire another shot. A semi-automatic rifle, by contrast, will automatically eject a fired cartridge and chamber a new cartridge, providing a user with one shot per trigger pull. would be too slow to operate and use for self-defense. In this instance a semi-automatic rifle could be necessary to defend against an attacker who is armed with a semi-automatic long gun, thus serving an important need under section 5(b). Moreover, unlike the unfettered discretion that the “may-issue” regimes discussed in Bruen allowed for,223N.Y. State Rifle & Pistol Ass’n v. Bruen, 142 S. Ct. 2111, 2156 (2022). section 5(c) severely limits the discretion of the designated firearms officer by requiring “an objectively reasonable basis based on clear and convincing evidence” to support a denial.224See supra Section III.B.

    Individuals could also obtain a license to possess a restricted firearm for dedicated hunting or sports shooting under section 9 of the Model Act. This is intended for individuals who regularly engage in hunting or sports shooting activities such as competitive shooting. This provision serves the purpose of limiting access to such particularly dangerous firearms as semi-automatic rifles while still allowing individuals to continue to engage in hunting, target practice, and shooting competitions using other kinds of firearms. The requirements that a person be regularly engaged in hunting or sports shooting and belong to an accredited hunting or sports shooting association is meant to keep semi-automatic rifles from being available to any adult for any purpose.

    At first blush, this restriction on semi-automatic rifles seems to violate the historical test created in Bruen,225Bruen, 142 S. Ct. at 2126. unless a historical analogue can be found. However, a close reading of Bruen’s test shows that, in reviewing section 9 of the Model Act, the burden to find a historical analogue would never shift to the government. The Bruen test states that the Second Amendment presumptively protects an individual’s conduct only when its “plain text covers an individual’s conduct.”226Id. The Court has held that the plain text of the Amendment covers “ ‘the individual right to possess and carry weapons in case of confrontation’ that does not depend on service in the militia.”227Id. at 2127 (quoting District of Columbia v. Heller, 554 U.S. 570, 592 (2008)). Section 9 of the Model Act, however, does not burden the “individual right to possess and carry weapons in case of confrontation”228Heller, 554 U.S. at 592. in the slightest. It restricts only the sporting use of certain weapons, not their self-defense use.229It is important to note that the Heller Court mentioned a right to hunting. Id. at 599 (“[M]ost [Americans] undoubtedly thought it even more important for self-defense and hunting.”). The Model Act accounts for this by allowing for permissive licensing for sporting purposes or hunting under section 8.

    The restriction on the sporting use of certain especially dangerous arms notwithstanding, individuals who wish to own a restricted firearm for self-defense have that option, subject only to a showing that the restricted firearm they wish to possess serves an important purpose that a non-restricted firearm cannot. This provision requiring an applicant to show that a restricted firearm serves an important purpose is also likely to be found constitutional under Bruen. Because section 5(a) concerns a restriction on firearms ownership for self-defense—unlike the sporting use contemplated in section 9—the “plain text” of the Second Amendment presumptively covers the conduct in question. Therefore, the burden would shift to the government to prove that section 5 burdens the right to bear arms for self-defense in a similar way and for a similar reason as a historical analogue to that regulation.230Bruen, 142 S. Ct. at 2132–33. The clear historical analogue for section 5 is the English prohibition on going armed with dangerous or unusual weapons.

    The offense of riding or going armed with dangerous or unusual weapons is a crime against the public peace, by terrifying the good people of the land, and is particularly prohibited by the statute of Northampton, 2 Edward III, c. 3 (Wearing Arms, 1328), upon pain of forfeiture of the arms, and imprisonment during the king’s pleasure . . . .2314 William Blackstone, Commentaries on the Laws of England *149 (William Carey Jones ed., Claitor’s Publ’g Div. 1976) (1765).

    Laws prohibiting going armed with dangerous or unusual weapons form a “long, unbroken line of common-law precedent”232Bruen, 142 S. Ct. at 2136. that was recognized in the United States following the adoption of the Second Amendment.233See Blackstone, supra note 231; 1 William Hawkins, Treatise of the Pleas of the Crown 489 (John Curwood ed., 8th ed. 1824) (“[P]ersons of quality are in no danger of offending against this statute [prohibiting affrays] by wearing common weapons . . . .”); State v. Langford, 10 N.C. (3 Hawks) 381, 383 (1824) (“[T]here may be an affray when there is no actual violence: as when a man arms himself with dangerous and unusual weapons . . . .”); State v. Huntly, 25 N.C. (3 Ired.) 418, 420 (1843) (“[T]he offence of riding or going about armed with unusual and dangerous weapons, to the terror of the people, was created by the statute . . . .”); State v. Lenier, 71 N.C. 288, 289 (1874); English v. State, 35 Tex. 473, 473 (1871) (rejecting Second Amendment challenge to law regulating the carrying of pistols, dirks, bowie knives, and other deadly weapons). Although the Court stated in Bruen that the English and colonial laws prohibiting affrays were not sufficiently analogous to New York’s proper cause requirement,234Bruen, 142 S. Ct. at 2143. it did not foreclose reliance on these laws to justify other firearms regulations.235Id. The Court went so far as to state that “colonial legislatures sometimes prohibited the carrying of ‘dangerous and unusual weapons’—a fact we already acknowledged in Heller.”236Id. (quoting District of Columbia v. Heller, 554 U.S. 570, 627 (2008)). Crucially, sections 4 and 5 of the Model Act concern dangerous and unusual weapons, not a restrictive carry licensing scheme like the one the respondents sought to justify in Bruen.

    In comparing sections 4 and 5 of the Model Act with the historical analogue, it is first necessary to establish whether the arms described in section 5(a) can fairly be described as “dangerous and unusual.” Clearly, compared with the types of arms the English law prohibited at the time of Blackstone’s Commentaries, the arms described in section 5(a) are extraordinarily dangerous and unusual. Even compared with other modern firearms, however, semi-automatic rifles that can be reloaded quickly are uniquely destructive.237Klay, supra note 218, at 152–54. They are capable of inflicting an incredible amount of damage in a short period of time, making them especially dangerous and unusual by any standard.238Id. The federal government acknowledged as much in making the transfer or possession of assault weapons unlawful. Public Safety and Recreational Firearms Use Protection Act, H.R. 4296, 103d Cong. (1994) (repealed 2004). The act was allowed to expire in 2004 in accordance with its sunset provision. Comparing the relative burdens of the historical prohibition and the Model Act, the latter clearly burdens the right to bear arms to a lesser degree than the historical analogue. The historical offense outright prohibits going armed with dangerous or unusual weapons while sections 4 and 5 only limit it to certain uses. Within the terms of the Model Act, these uses include the “law-abiding citizen’s right to armed self-defense.”239Bruen, 142 S. Ct. at 2133. Moreover, the historical analogue and the Model Act burden the right for a similar purpose—to prevent especially dangerous and frightening arms from being widespread and to prevent individuals from terrorizing others with these arms.240Blackstone, supra note 231.

    C.  Public Carry

    The portion of the Model Act regulating public carry of firearms for self-defense—sections 6 and 7—is perhaps as restrictive as courts will allow under Bruen. Of course, sections 6(b) and 7(b) make clear that the Model Act establishes a shall-issue public carry licensing regime, as required by Bruen.241Bruen, 142 S. Ct. at 2156. Sections 6 and 7 also require an applicant to have a valid eligibility certificate, which is where most of the requirements that help ensure safe use of a firearm are listed. The eligibility certificate requirements would not be likely to face much resistance from the courts because they do not burden the right per se;242See supra Section IV.A. rather, they are “designed to ensure only that those bearing arms in the jurisdiction are, in fact, ‘law-abiding, responsible citizens.’ ”243Bruen, 142 S. Ct. at 2138 n.9 (quoting District of Columbia v. Heller, 554 U.S. 570, 635 (2008)). The Model Act leaves to the states the decision of whether one type of carry—open or concealed—should be banned.244Id. at 2150.

    CONCLUSION

    The governing case law concerning the Second Amendment greatly limits how states can restrict firearm ownership. The Supreme Court’s historical approach to Second Amendment challenges places many regulations many people find desirable outside the realm of constitutionality. This does not mean, however, that all reasonable regulations are impossible to implement. The Model Firearms Control Act presented in this Note is an initial step toward a comprehensive firearms licensing system that can serve to keep Americans safer while respecting their right to armed self-defense. The relatively limited nature of the regulations advocated in this Note is simply an acknowledgement of the reality that the Supreme Court has adopted a sweeping interpretation of the Second Amendment irrespective of its true meaning, which may best be left to historians. Given its current membership, the Supreme Court will not, for the foreseeable future, overturn its trilogy of Second Amendment precedents, so all states can do is implement the safest gun control solutions that governing case law will allow. On a later day, perhaps, a differently constituted Supreme Court will reconsider Bruen and replace its untenable historical inquiry with some form of means-end scrutiny. If and when that happens, the Model Act proposed in this Note can be expanded to be more restrictive while still respecting the individual right to armed self-defense. This Note offers an initial large step in implementing such solutions.

97 S. Cal. L. Rev. 211

Download

* Senior Editor, Southern California Law Review, Volume 97; J.D. Candidate 2024, University of Southern California Gould School of Law; B.S. Criminology & Criminal Justice 2020, California State University, Long Beach. Thank you to Professor David B. Cruz for his invaluable feedback and to the Southern California Law Review for their meticulous editing. A very special thank you to Alyssa Polito whose unwavering support during my time in law school has made this Note possible.

“Shelby County” to Clean Air Act: Evaluating the Constitutionality of California’s Clean Air Act Waiver Under the Equal Sovereignty Principle

In August 2022, California promulgated the Advanced Clean Cars II regulation, banning all sales of new gasoline-powered cars in the state by 2035. Transportation is the largest source of air pollution in California, responsible for nearly 40% of greenhouse gas (“GHG”) emissions; thus, the regulation is a crucial step towards meeting the state’s carbon neutrality and climate goals. California has the unique authority to regulate motor vehicle emissions due to a waiver exemption in the Clean Air Act. Congress recognized California’s expertise and unique air pollution challenges early on by authorizing just two standards: the national and California standards. Over the last five decades, California has received over one hundred waivers for its motor vehicle emission standards. However, in May 2022, seventeen states challenged the constitutionality of the waiver provision in Ohio v. EPA (pending in the D.C. Circuit Court of Appeals as of the publication of this Note), alleging, inter alia, that it violates the equal sovereignty principle—the idea that states must have equal political authority—by allowing only California to set new vehicle emission standards. These states further argue that California cannot regulate GHGs because climate change is a global problem not unique to California. To date, no court has addressed the constitutionality of the Clean Air Act under the equal sovereignty principle. Thus, this Note takes the principle seriously and analyzes how courts historically have applied it. In 2013, the Supreme Court developed the equal sovereignty principle as a meaningful concept for the first (and last) time in Shelby County v. Holder to invalidate section 4 of the Voting Rights Act. This Note applies the test established in Shelby County to the Clean Air Act waiver at issue, arguing that the equal sovereignty principle does not apply to the Clean Air Act, and even if it were to apply, the Clean Air Act waiver provision remains constitutional because Congress’s reasons for granting California an exemption remain relevant. California’s pioneering role in early air pollution control, its large economy, and its characteristic geographic and climate conditions put the state in a unique position to protect public health by regulating automobile emissions, while the state faces increasingly formidable threats from climate change that have exacerbated the local air pollution problems that initially compelled its motor vehicle regulations. Thus, even as California’s motor vehicle regulations have shifted from reducing local smog to reducing GHG emissions, California’s current needs continue to justify its differential treatment—maintaining, and perhaps even strengthening, the Clean Air Act waiver’s relevance in the twenty-first century.

INTRODUCTION

In August 2022, the California Air Resources Board (“CARB”), California’s chief air pollution regulator, promulgated the Advanced Clean Cars II regulation, which bans the sale of new gasoline-powered cars in California by 2035.1Advanced Clean Cars II Regulations: All New Passenger Vehicles Sold in California to be Zero Emissions by 2035, Cal. Air Res. Bd., https://ww2.arb.ca.gov/our-work/programs/advanced-clean-cars-program/advanced-clean-cars-ii [https://perma.cc/A9WT-T2BP]; Cal. Code Regs. tit. 13, § 1962.4 (2022). Transportation is the largest source of air pollution in the state, responsible for nearly 40% of greenhouse gas (“GHG”) emissions, 80% of nitrogen oxide pollution, and 90% of diesel particulate matter pollution.2Current California GHG Emission Inventory Data, Cal. Air Res. Bd., https://ww2.arb.ca.gov/ghg-inventory-data [https://perma.cc/L9KM-VCG3]; Transforming Transportation, Cal. Energy Comm’n, https://www.energy.ca.gov/about/core-responsibility-fact-sheets/transforming-transportation [http://perma.cc/LAS2-MAYL]. CARB estimates that the new rule will result in significant climate, economic, and public health benefits. By 2040, the regulation is projected to result in a 50% reduction in GHG emissions from cars, pickup trucks, and SUVs.3California Moves to Accelerate to 100% New Zero-Emission Vehicle Sales by 2035, Cal. Air Res. Bd. (Aug. 25, 2022), https://ww2.arb.ca.gov/news/california-moves-accelerate-100-new-zero-emission-vehicle-sales-2035 [https://perma.cc/5GRX-9NXR]. Taking gas cars off the road would eliminate the equivalent of 395 million metric tons of carbon dioxide emissions, which is analogous to avoiding the combustion of 915 million barrels of petroleum or shutting down more than one hundred coal plants for a year.4Id. From 2026 to 2040, the decrease in pollution should lead to 1,290 fewer cardiopulmonary deaths, 460 fewer hospital admissions for cardiovascular or respiratory illness, and 650 fewer emergency room visits for asthma.5Id. Thus, the regulation is a crucial step towards meeting the state’s carbon neutrality and climate goals.6Id. (“The ACC II regulation is a major tool in the effort to reach the SB 32 target of reducing greenhouse gases an additional 40% below 1990 levels by 2030 . . . . Ending sales of vehicles powered by fossil fuels is a critical element in the state’s efforts to achieve carbon neutrality by 2045 or sooner.”).

The regulations that California enacts are hugely influential; thus, the implications of California’s ability to implement motor vehicle regulations are extensive. If California were a country, it would be the tenth largest auto market in the world.7Naveena Sadasivam, It’s Official: California is Phasing Out Gas-Powered Cars by 2035, Grist (Aug. 25, 2022), https://grist.org/transportation/california-gas-car-ban-electric-vehicles [https://perma.cc/2XPY-J5HH]. As of May 13, 2022, seventeen states and the District of Columbia have adopted California’s Low-Emission Vehicle (“LEV”) and Zero-Emission Vehicle (“ZEV”) regulations under section 177 of the Clean Air Act, which allows other states to adopt California’s approved standards instead of the federal standards.8States That Have Adopted California’s Vehicle Standards Under Section 177 of the Federal Clean Air Act, Cal. Air Res. Bd. (May 13, 2022), https://ww2.arb.ca.gov/sites/default/files/2022-05/%C2%A7177_states_05132022_NADA_sales_r2_ac.pdf [https://perma.cc/EM9D-QLM9]. California alone makes up 11% of U.S. new light-duty vehicle sales, or 40.1% when combined with the states that have already adopted its rules.9Id. New York was the second state to ban sales of gas-powered cars by 2035 as part of its plan to increase EV adoption.10Kira Bindrim, NY Implements 2035 All-EV Plan After California Clears the Way, Bloomberg (Sept. 29, 2022, 1:57 PM), https://www.bloomberg.com/news/articles/2022-09-29/new-york-follows-california-in-banning-sale-of-gas-cars-by-2035 [https://perma.cc/N7N6-LUSS]. In February 2021, New York passed a law requiring all new passenger cars and trucks sold in the state to produce zero emissions by 2035,11Assemb. B. 4302, 2021–2022 Leg. Reg. Sess. (N.Y. 2021). and in September 2022, after California finalized its own ban, New York followed California in requiring all new vehicles sold by 2035 to be zero-emissions, setting in motion the regulatory process to implement the law.12Press Release, Kathy Hochul, Governor of the State of New York, Governor Hochul Drives Forward New York’s Transition to Clean Transportation (Sept. 29, 2022), https://www.governor.ny.gov/news/governor-hochul-drives-forward-new-yorks-transition-clean-transportation [https://perma.cc/8EJ3-NPTG]. In August 2022, Massachusetts Governor Charlie Baker also signed climate change legislation to end new sales of gas-powered cars in the state by 2035.13Keith Goldberg, Calif. Sews Up Regs to End Gas Car Sales by 2035, Law360 (Aug. 25, 2022, 6:52 PM), https://www.law360.com/articles/1524638/calif-sews-up-regs-to-end-gas-car-sales-by-2035 [https://perma.cc/RQK3-HTU3].

California has the unique ability to implement motor vehicle emissions regulations because of an exception in the Clean Air Act.1442 U.S.C. § 7543(b)(1). While the Clean Air Act generally prohibits states from setting vehicle emission standards,1542 U.S.C. § 7543(a). it provides a waiver exemption under section 209(b)(1) that allows California to set more stringent vehicle emission standards than the federal government.1642 U.S.C. § 7543(b)(1). While the waiver does not reference California by name, it was clearly intended for California because California was the only state that met the requirement of adopting motor vehicle emission standards prior to March 30, 1966. H.R. Rep. No. 90-728, at 49 (1967). Given California’s pioneering role in motor vehicle regulations and unique air pollution problems, Congress recognized California’s expertise early on in the history of federal air pollution regulation.17See Air Quality Act of 1967, S. Rep. No. 90-403, at 33 (“California’s unique problems and pioneering efforts justified a waiver . . . . [I]n the 15 years that auto emission standards have been debated and discussed, only the State of California has demonstrated compelling and extraordinary circumstances sufficiently different from the Nation as a whole . . . .”). However, in May 2022, seventeen Republican-led states filed a lawsuit, Ohio v. EPA, challenging California’s ability to set its own pollution rules and demanding that the U.S. Environmental Protection Agency (“EPA”) revoke the waiver.18Petition for Review, Ohio v. EPA, No. 22-1081 (D.C. Cir. May 12, 2022). The petitioner states claimed that the waiver provision is unconstitutional because it violates the so-called equal sovereignty principle—the idea that states must have equal political authority—by only empowering California to set new vehicle emission standards.19See Brief for Petitioners at 28, Ohio v. EPA, No. 22-1081 (D.C. Cir. Nov. 11, 2022). The petitioners additionally argued that Congress cannot allow California alone to regulate climate change, which is a global problem not unique to California.20Id. at 13. Because California has shifted from regulations to reduce smog and local air pollution to GHG regulations to address global climate change, the petitioners essentially argued that circumstances have changed enough since Congress enacted the waiver provision in 1967 that California’s special treatment is no longer justified.21See id. at 13, 30–33.

This Note takes the equal sovereignty claim seriously and argues that the Clean Air Act waiver provision remains constitutional under the equal sovereignty principle. Part I provides relevant background on the waiver provision and history of California’s waiver requests. It then summarizes the equal sovereignty principle arguments made in the pending Ohio v. EPA lawsuit and provides relevant history on how courts have applied the principle leading up to Shelby County v. Holder,22Shelby County v. Holder, 570 U.S. 529, 540 (2013). the first time the Supreme Court held a statute unconstitutional based on the equal sovereignty principle. Part II argues that the equal sovereignty principle does not apply to the Clean Air Act, but even if it were to apply, the test from Shelby County does not invalidate the Clean Air Act waiver provision. This Note concludes by offering final thoughts on the equal sovereignty claim and underscoring the implications of Ohio v. EPA in California’s ability to continue to lead the nation in addressing GHG emissions.

I.  BACKGROUND

A.  The Clean Air Act and EPA Waiver Provision

California’s ability to implement its own motor vehicle standards stems from the Clean Air Act. Congress passed the Clean Air Act in response to air pollution crises in the mid-20th century resulting from industrialization.23Clean Air Act Requirements and History, EPA, https://www.epa.gov/clean-air-act-overview/clean-air-act-requirements-and-history [https://perma.cc/HL9S-DUXJ]. “Killer fog” events, where a deadly mix of pollution and fog covered cities in the United States and around the world, spurred federal regulation of air pollution.24The 1948 Donora, Pennsylvania killer fog killed at least 20 people and left 5,900 ill. Lorraine Boissoneault, The Deadly Donora Smog of 1948 Spurred Environmental Protection—But Have We Forgotten the Lesson?, Smithsonian (Oct. 26, 2018), https://www.smithsonianmag.com/history/deadly-donora-smog-1948-spurred-environmental-protection-have-we-forgotten-lesson-180970533 [https://perma.cc/QXH6-BJ4N]; Elizabeth T. Jacobs, Jefferey L. Burgess & Mark B. Abbott, The Donora Smog Revisited: 70 Years After the Event That Inspired the Clean Air Act, 108 Am. J. Pub. Health S2, S85–S88 (2018). The 1952 London Killer Fog killed between 8,000 and 12,000 people. Christopher Klein, When the Great Smog Smothered London, History (Dec. 6, 2012), https://www.history.com/news/the-killer-fog-that-blanketed-london-60-years-ago [https://perma.cc/BS36-3M7Z]. In 1955, Congress enacted the Air Pollution Control Act, the first national air pollution legislation.25Air Pollution Control Act of 1955, Pub. L. No. 84-159, 69 Stat. 322, 322. Continuing “killer fog” incidents in the United States then prompted Congress to pass the 1963 Clean Air Act, which established grant and research programs to support states in their air pollution control efforts but left air pollution regulation primarily to the states.26Clean Air Act of 1963, Pub. L. No. 88-206, 77 Stat. 392, 393.

California was the first state to regulate emissions from cars.27History, Cal. Air Res. Bd., https://ww2.arb.ca.gov/about/history [https://perma.cc/BA4F-FJXN]. The first recognized episodes of smog occurred in Los Angeles in 1943, and in the 1950s, a California researcher determined that the automobile was the main cause of the smog.28Id.; Timeline of Major Accomplishments in Transportation, Air Pollution, and Climate Change, EPA, https://www.epa.gov/transportation-air-pollution-and-climate-change/timeline-major-accomplishments-transportation-air [https://perma.cc/ZS88-ZEXJ]. In 1966, California established the first tailpipe emissions standards in the nation.29Cal. Air Res. Bd., supra note 27.

Congress continued to enact new statutes in response to California’s regulations.30The 1967 Air Quality Act regulations for controlling motor vehicle emissions “were patterned after those . . . in effect in California.” 113 Cong. Rec. S32478 (daily ed. Nov. 14, 1967) (remarks by Sen. George Murphy of California). The 1967 Air Quality Act amended the 1963 Clean Air Act, moving towards a uniform federal policy by requiring national air quality criteria, which states would then implement.31Air Quality Act of 1967, Pub. L. No. 90-148, 81 Stat. 485, 485–86. It was also the first statute to give preemptive power to the federal government to adopt and enforce standards relating to the control of emissions from new motor vehicles.32Id. at 501. However, Congress added a waiver provision exempting California from the preemption provision when California could demonstrate a need for more stringent standards than those the EPA established.33“The Secretary shall . . . waive application of [federal preemption] . . . to any State which has adopted standards . . . for the control of emissions from new motor vehicles or new motor vehicle engines prior to March 30, 1966, unless he finds that such State does not require standards more stringent than applicable Federal standards to meet compelling and extraordinary conditions . . . .” Air Quality Act of 1967, Pub. L. No. 90-148, § 208(b), 81 Stat. 485, 501. While the waiver does not reference California by name, it was clearly intended for California because California was the only state that met the requirement of adopting motor vehicle emission standards prior to March 30, 1966.34H.R. Rep. No. 90-728, at 49 (1967). Thus, Congress acknowledged California’s expertise early on in the history of federal air pollution regulation.

In fact, the Clean Air Act is a paradigmatic example of cooperative federalism, under which “States and the Federal Government [are] partners in the struggle against air pollution.”35Gen. Motors Corp. v. United States, 496 U.S. 530, 532 (1990). The federal preemption provision reflects Congress’s interest in allowing automobile manufacturers to produce uniform automobiles for a national market and benefit from the economies of large-scale production without having to accommodate multiple state standards.36H.R. Rep. No. 90-728, at 21 (1967); see also id. at 50. Congress acknowledged the complex nature of automobile manufacturing and noted the importance of ensuring that automobile manufacturers obtain “clear and consistent answers” concerning emission standards.37Id. at 21. Courts have also interpreted that Congress preempted the field of vehicle emission regulation “to ensure uniformity throughout the nation, and to avoid the undue burden on motor vehicle manufacturers which would result from different state standards.”38Motor Vehicle Mfrs. Ass’n v. New York State Dep’t. of Env’t. Conservation, 810 F. Supp. 1331, 1337 (N.D.N.Y. 1993), aff’d in part, rev’d in part, 17 F.3d 521 (2d Cir. 1994). However, given California’s lead in early motor vehicle regulations and Congress’s additional interest in having California as a “laboratory for innovation,”39Motor & Equip. Mfrs. Ass’n v. EPA (MEMA I), 627 F.2d 1095, 1111 (D.C. Cir. 1979). Congress intentionally struck a balance between having one national standard and fifty different state standards by authorizing just two standards, the national and California standards.40See S. Rep. No. 90-403, at 33 (1967) (“California’s unique problems and pioneering efforts justified a waiver . . . .[I]n the 15 years that auto emission standards have been debated and discussed, only the State of California has demonstrated compelling and extraordinary circumstances sufficiently different from the Nation as a whole . . . .”); 113 Cong. Rec. H30975 (daily ed. Nov. 2, 1967) (remarks by Rep. John Moss) (“[The California waiver] permits California to continue a role of leadership which it has occupied among the States of this Union for at least the last two decades . . . . [I]t offers a unique laboratory, with all of the resources necessary, to develop effective control devices which can become a part of the resources of this Nation and contribute significantly to the lessening of the growing problems of air pollution throughout the Nation.”); see also Engine Mfrs. Ass’n v. EPA, 88 F.3d 1075, 1080 (D.C. Cir. 1996) (“Rather than being faced with 51 different standards, as they had feared, or with only one, as they had sought, manufacturers must cope with two regulatory standards . . . .”); Motor & Equip. Mfrs. Ass’n v. Nichols (MEMA II), 142 F.3d 449, 463 (D.C. Cir. 1998). This balance allowed California to continue to innovate and improve its air quality without creating a practical nightmare for automakers and interstate commerce.41Members of Congress favored states’ rights, but also were concerned that having 50 different sets of requirements related to emissions controls would “unduly burden interstate commerce.” H.R. Rep. No. 95-294, at 309 (1977).

The 1970 Clean Air Act Amendments, which form the basis of the contemporary federal Clean Air Act, authorized the development of federal and state regulations to limit emissions from stationary (industrial) and mobile sources (including automobiles).42Clean Air Act Amendments of 1970, Pub. L. No. 91-604, 84 Stat. 1676, 1678; Evolution of the Clean Air Act, EPA, https://www.epa.gov/clean-air-act-overview/evolution-clean-air-act [https://perma.cc/7XMF-6QVB]. Section 109 requires the EPA Administrator to establish basic requirements for ambient air quality, known as National Ambient Air Quality Standards (“NAAQS”), for particular criteria pollutants, which the states would be required to meet.43Clean Air Act Amendments of 1970, Pub. L. No. 91-604, § 109(a)(1), § 110(a)(1), 84 Stat. 1676, 1679–80. The current list of criteria pollutants includes sulfur dioxide, particulate matter, nitrogen oxide, carbon monoxide, ozone, and lead, but does not include carbon dioxide.44Criteria Air Pollutants, EPA, https://www.epa.gov/criteria-air-pollutants [https://perma.cc/Y9JR-T8K6].

In 1977, Congress revised the provision to read as it does today. Section 202(a)(1) requires the EPA Administrator to establish motor vehicle emissions standards for pollutants “which, in his judgment, cause or contribute to air pollution which may reasonably be anticipated to endanger public health or welfare.”45Clean Air Act Amendments of 1977, Pub. L. No. 95-95, § 401(d)(1), 91 Stat. 685, 791. The 1977 Clean Air Act Amendments strengthened the deference given to California under the waiver provision in two significant ways. First, the 1977 Amendments revised section 209(b)(1) by requiring the EPA Administrator to grant a preemption waiver for California “if the State determines that the State standards will be, in the aggregate, at least as protective of public health and welfare as applicable Federal standards.”46Clean Air Act Amendments of 1977, Pub. L. No. 95-95, § 207, 91 Stat. 685, 755 (emphasis added). This amendment allows California, rather than the EPA, to make its own determination as to whether the regulations are sufficiently protective of public health and welfare. It also allows California to make this determination by looking at the entire program as a whole, rather than evaluating each regulation individually. Thus, as long as the entire set of regulations is more protective than the federal system, the EPA must allow California to implement these measures. The EPA Administrator can deny the waiver only if the state’s determination is “arbitrary and capricious” or the state does not need its standards to meet “compelling and extraordinary conditions.”47Id. Second, the 1977 Amendments added section 177, which enhanced the strength of California’s motor vehicle emissions regulations by allowing other states to adopt California’s approved standards in lieu of the federal standards.48Clean Air Act Amendments of 1977, Pub. L. No. 95-95, § 177, 91 Stat. 685, 750. According to the House Report, the Committee on Interstate and Foreign Commerce makes clear that it sought to “ratify and strengthen the California waiver provision . . . to afford California the broadest possible discretion in selecting the best means to protect the health of its citizens and the public welfare.”49H.R. Rep. No. 95-294, at 301–02 (1977). The legislative and statutory history thus suggests that Congress intended to give California broad discretion to regulate air pollutants in the way it deems most appropriate to protect public health and welfare.

B.  History of California’s Motor Vehicle Regulations and Waiver Requests

The Clean Air Act section 209(b)(1) waiver reflects a five-decade history of allowing California to implement motor vehicle emissions standards that are more stringent than federal government standards.50Pollution Standards Authorized by the California Waiver: A Crucial Tool for Fighting Air Pollution Now and in the Future, Cal. Air Res. Bd. (Sept. 17, 2019), https://ww2.arb.ca.gov/resources/fact-sheets/pollution-standards-authorized-california-waiver-crucial-tool-fighting-air [https://perma.cc/P6EX-HUGH]; Emily Wimberger & Hannah Pitt, Come and Take It: Revoking the California Waiver, Rhodium Grp. (Oct. 28, 2019), https://rhg.com/research/come-and-take-it-revoking-the-california-waiver [https://perma.cc/3Q28-6RBA] (“Since 1970, the federal government has granted California over 100 waivers . . . .”); see Vehicle Emissions California Waivers and Authorizations, EPA, https://www.epa.gov/state-and-local-transportation/vehicle-emissions-california-waivers-and-authorizations [https://perma.cc/VA5H-RSVG] (documenting all the waivers the EPA has granted). California was granted its first waiver in 1968 and has since received over one hundred waivers for a range of new or amended motor vehicle and motor vehicle engine standards.51Pollution Standards Authorized by the California Waiver: A Crucial Tool for Fighting Air Pollution Now and in the Future, Cal. Air Res. Bd. (Sept. 17, 2019), https://ww2.arb.ca.gov/resources/fact-sheets/pollution-standards-authorized-california-waiver-crucial-tool-fighting-air [https://perma.cc/P6EX-HUGH]; Vehicle Emissions California Waivers and Authorizations, EPA, https://www.epa.gov/state-and-local-transportation/vehicle-emissions-california-waivers-and-authorizations [https://perma.cc/VA5H-RSVG] (documenting all the waivers the EPA has granted). Smog in Los Angeles initially spurred California to adopt statewide standards to regulate criteria pollutants,52See infra Section I.A. and CARB has consistently developed the first emission standards in the nation.53The California Air Resources Board (“CARB”) developed the nation’s first tailpipe emissions standards for hydrocarbons and carbon monoxide in 1966, oxides of nitrogen in 1971, and particulate matter from diesel-fueled vehicles in 1982, as well as catalytic converters in the 1970s. More recently, CARB has delved into regulations seeking to mitigate climate change by encouraging Low-Emission Vehicles (“LEVs”). It promulgated LEV regulations that established criteria pollutant regulations for light and medium-duty vehicles in 1990 for the 1994–2003 model years (LEV I), and in 1999 for the 2004 model year and after (LEV II). Low-Emission Vehicle Program, Cal. Air Res. Bd., https://ww2.arb.ca.gov/our-work/programs/low-emission-vehicle-program/about [https://perma.cc/R7KV-ME7L]; Low-Emission Vehicle (LEV II) Program, Cal. Air Res. Bd., https://ww2.arb.ca.gov/our-work/programs/advanced-clean-cars-program/lev-program/low-emission-vehicle-lev-ii-program [https://perma.cc/MG4U-3U6M].

As California transitioned from regulating criteria pollutants to promulgating regulations that address GHG emissions, certain EPA administrations began to challenge its waiver requests, leading to the ping-ponging back and forth between administrations. In 2002, recognizing that global warming would impose “compelling and extraordinary impacts” on California, the state enacted Assembly Bill (AB) 1493, Chapter 200.54Assemb. B. 1493, Ch. 200, 2001–2002 Leg. Reg. Sess. (Cal. 2002). The bill acknowledged that motor vehicle emissions are a major source of the state’s GHG emissions and that reducing GHG emissions is critical to slowing down the effects of global warming and protecting public health and the environment.55Id. The bill directed CARB to adopt regulations that achieve the “maximum feasible . . . reduction of greenhouse gas emissions” from passenger vehicles, beginning with the 2009 model year.56Id. Thus, in 2004, CARB approved the first regulations in the nation that control GHG emissions from motor vehicles (Pavley regulations), which applied to new vehicles for the 2009–2016 model years.57Low-Emission Vehicle Program, Cal. Air Res. Bd., supra note 53.

In December 2005, CARB requested a waiver to allow California to enforce its new GHG emission standards.58California’s Greenhouse Gas Vehicle Emission Standards Under Assembly Bill 1493 of 2002 (Pavley), Cal. Air Res. Bd., https://ww2.arb.ca.gov/californias-greenhouse-gas-vehicle-emission-standards-under-assembly-bill-1493-2002-pavley [https://perma.cc/6T52-5YNF]. The EPA delayed action pending the outcome of litigation regarding whether the EPA had authority to regulate GHG emissions under the Clean Air Act, as the Clean Air Act did not explicitly regulate GHG emissions at the time.59Letter from John B. Stephenson, Director, Natural Resources and Environment, to Congressional Requesters (Jan. 16, 2009) (on file with the United States Government Accountability Office). The Supreme Court addressed GHG emissions for the first time in Massachusetts v. EPA, holding in a 5-4 decision that carbon dioxide is considered an “air pollutant” that the EPA may regulate under section 202(a)(1) of the Clean Air Act.60Massachusetts v. EPA, 549 U.S. 497, 532 (2007). Thus, the Court held that the EPA has the statutory authority to regulate GHG emissions from new motor vehicles and that Congress provided the EPA with the flexibility to address new air pollutant threats that the EPA determines endanger the public welfare.61Id.

Despite the Supreme Court ruling, in March 2008, the Bush administration’s EPA denied the waiver for the Pavley regulations, which was the first time the EPA denied a waiver for California.62California State Motor Vehicle Pollution Control Standards, Notice of Decision Denying a Waiver of Clean Air Act Preemption, 73 Fed. Reg. 12156, 12157 (Mar. 6, 2008) [hereinafter 2008 Waiver Denial]. In its decision, the EPA deviated from the traditional interpretation of the “compelling and extraordinary” waiver criteria6342 U.S.C. § 7543(b)(1); see Rachel L. Chanin, California’s Authority to Regulate Mobile Source Greenhouse Gas Emissions, 58 N.Y.U. Ann. Surv. Am. L. 699, 723 (2001); California State Motor Vehicle Pollution Control Standards, 49 Fed. Reg. 18887, 18889–92 (May 3, 1984). to narrowly interpret that Congress authorized the EPA to grant a waiver only when “California standards were necessary to address peculiar local air quality problems,” as opposed to global climate change problems.642008 Waiver Denial, 73 Fed. Reg. at 12161. Unlike California’s previous motor vehicle programs, which addressed local smog problems, the GHG emission standards aimed to address climate change. Thus, the EPA determined that California did not need its new motor vehicle standards to meet “compelling and extraordinary” conditions related to GHG emissions because emissions from California cars “become one part of the global pool of GHG emissions”65Id. at 12160. and do not directly cause elevated concentrations of GHGs in the region.66Id. at 12162 (“The local climate and topography in California have no significant impact on the long-term atmospheric concentrations of greenhouse gases in California.”). Alternatively, the EPA determined that because climate change is a global issue, the impacts of climate change in California were not sufficiently unique and different.67Id. at 12168.

In July 2009, the Obama administration’s EPA reversed the 2008 denial and granted California’s waiver request to enforce its GHG emission standards for model year 2009 and later new motor vehicles.68Notice of Decision Granting a Waiver of Clean Air Act Preemption, 74 Fed. Reg. 32744, 32746 (July 8, 2009) [hereinafter 2009 Waiver Grant]. As the EPA stated, CARB has repeatedly demonstrated the need for its motor vehicle program to address “compelling and extraordinary” conditions in California, and Congress did not intend to allow California to address only local or regional air pollution problems.69Id. at 32761. Rather, Congress intended California to have broad discretion and autonomy, acting as a pioneer and a “laboratory for innovation.”70Id. (citing Motor & Equip. Mfrs. Ass’n v. EPA (MEMA I), 627 F.2d 1095, 1111 (D.C. Cir. 1979)); see S. Rep. No. 90-403, at 33 (1967) (“The Nation will have the benefit of California’s experience with lower standards which will require new control systems and design. In fact California will continue to be the testing area for such lower standards and should those efforts to achieve lower emission levels be successful it is expected that the Secretary will . . . give serious consideration to strengthening the Federal standards.”). Thus, narrowing the waiver’s scope would hinder California from implementing motor vehicle programs “as it deems appropriate to protect the health and welfare of its citizens.”712009 Waiver Grant, 74 Fed. Reg. at 32761. In contrast to the 2008 EPA’s reasoning, the 2009 EPA determined that the impacts of global climate change can exacerbate the local air pollution problem.72Id. at 32763. It found compelling California’s assessment that its GHG standards are linked to improving California’s smog problems and that higher temperatures from global warming will exacerbate California’s high ozone levels and the “climate, topography, and population factors conducive to smog formation in California, which were the driving forces behind Congress’s inclusion of the waiver provision in the Clean Air Act.”73Id. The EPA noted that California’s GHG regulations will reduce greenhouse gas concentrations, even if only slightly, and “every small reduction is helpful . . . .”74Id. at 32766. Given California’s unique geographical and climatic conditions that foster extreme air quality issues, its ongoing need for dramatic emissions reductions, and growth in its vehicle population and use, the EPA determined that California’s need met “compelling and extraordinary” conditions.75Id. at 32760. Still, the EPA acknowledged that “conditions in California may one day improve such that it no longer has the need for a separate motor vehicle program.”76Id. at 32762.

In 2012, CARB adopted the Advanced Clean Cars I (“ACC I”) regulations to increase the stringency of criteria pollutant and GHG emission standards for new passenger vehicles for the 2015–2025 model years.77The regulations consisted of two programs: (1) the Low Emission Vehicle program, designed for cars to emit 75% less smog-forming pollution (criteria pollutants) than the average car sold in 2012 and to reduce GHG emissions by about 40% from 2012 model year vehicles by 2025; and (2) the Zero Emission Vehicle program, which requires manufacturers to ensure that about 22% of their California sales consist of zero-emission vehicles and plug-in hybrids by 2025. Advanced Clean Cars Program, Cal. Air Res. Bd., https://ww2.arb.ca.gov/our-work/programs/advanced-clean-cars-program/about [https://perma.cc/W2R9-KFF7]. In 2013, the Obama administration’s EPA granted California a waiver for its ACC I regulations.78Notice of Decision Granting a Waiver of Clean Air Act Preemption, 78 Fed. Reg. 2112, 2145 (Jan. 9, 2013) [hereinafter 2013 Waiver Grant]. The EPA largely followed the 2009 waiver decision in determining that the new standards continued to meet “compelling and extraordinary” conditions.79Id. at 2131. The EPA found a rational connection between CARB’s emission standards and long-term air quality goals,80Id. (“Whether or not the ZEV standards achieve additional reductions by themselves above and beyond the LEV III GHG and criteria pollutant standards, the LEV III program overall does achieve such reductions, and EPA defers to California’s policy choice of the appropriate technology path to pursue to achieve these emissions reductions.”). The long-term goals were to have ZEVs be nearly 100% of new vehicle sales between 2040 and 2050, and reduce GHG emissions by 80% below 1990 levels by 2050. Id. at 2131–32. as well as compelling and extraordinary conditions within the state pertaining to the effects of pollution.81CARB noted: “Record-setting fires, deadly heat waves, destructive storm surges, loss of winter snowpack—California has experienced all of these in the past decade and will experience more in the coming decades . . . . In California, extreme events such as floods, heat waves, droughts and severe storms will increase in frequency and intensity. Many of these extreme events have the potential to dramatically affect human health and well-being, critical infrastructure and natural systems.” Id. at 2129.

In September 2019, in an unprecedented move, the Trump administration’s EPA revoked the 2013 waiver, marking the first time the EPA retroactively withdrew a previously granted waiver.82The Safer Affordable Fuel-Efficient (SAFE) Vehicles Rule Part One: One National Program, 84 Fed. Reg. 51310, 51310 (Sept. 27, 2019) [hereinafter 2019 Waiver Withdrawal]. The EPA and National Highway Traffic Safety Administration (NHTSA) issued a joint rulemaking that withdrew the waiver of California’s GHG and ZEV standards that were part of the ACC I program. The EPA went a step further than its 2008 waiver decision, narrowly interpreting that “Congress did not intend the waiver provision . . . to be applied to California measures that address pollution problems of a national or global nature,” but only conditions “extraordinary” with respect to California; that is, “with a particularized nexus to emissions in California and to topographical or other features peculiar to California.”83Id. at 51347. The EPA argued that climate change caused by carbon dioxide emissions is not a local air pollution problem and that California’s new motor vehicle standards deviated too far from what Congress intended in granting California a waiver.84Id. at 51350 n.285 (“Attempting to solve climate change, even in part, through the Section 209 waiver provision is fundamentally different from that section’s original purpose of addressing smog-related air quality problems.”) (quoting the SAFE proposal). The EPA concluded that California’s GHG standards were missing a specific connection to local features, and thus excluded GHG regulation from the scope of the waiver.85Id. at 51347, 51350.

In March 2022, the Biden administration’s EPA rescinded the 2019 waiver withdrawal, restoring the 2013 waiver and California’s authority to enforce its GHG emission standards and ZEV sales mandate.86Advanced Clean Car Program; Reconsideration of a Previous Withdrawal of a Waiver of Preemption; Notice of Decision, 87 Fed. Reg. 14332, 14332 (Mar. 14, 2022) [hereinafter 2022 Waiver Reconsideration]. In determining that California has a compelling need for its GHG standards and ZEV sales mandate, the EPA essentially reverted back to its 2013 analysis, maintaining that pollution continues to pose a distinct problem in California.87Id. at 14352–53, 14367. The EPA saw no reason to distinguish between local and global air pollutants, reasoning that all pollutants play a role in California’s local air quality problems and that the EPA should provide deference to California in its comprehensive policy choices for addressing them.88Id. at 14363. The 2022 EPA refuted the 2019 EPA’s premise that GHG emissions from motor vehicles in California do not pose a local air quality issue,89Id. at 14365–66. noting that criteria pollution and GHGs have interrelated and interconnected impacts on local air quality.90“[T]he Agency [in SAFE 1] failed to take proper account of the nature and magnitude of California’s serious air quality problems, including the interrelationship between criteria and GHG pollution.” Id. at 14334. “The air quality issues and pollutants addressed in the ACC program are interconnected in terms of the impacts of climate change on such local air quality concerns such as ozone exacerbation and climate effects on wildfires that affect local air quality.” Id. at 14334 n.10. CARB also attributed GHG emissions reductions to vehicles in California, projecting that the standards will reduce car CO2 emissions by about 4.9% a year. Id. at 14366.

Congress recently expanded the Clean Air Act to include GHGs, clarifying that GHGs are pollutants under the Clean Air Act. On August 16, 2022, President Biden signed the Inflation Reduction Act into law, the single largest climate package in U.S. history, which will invest almost $370 billion in clean energy and other climate-related measures over the next ten years, and is expected to reduce U.S. carbon emissions by 40% by 2030 compared to 2005 levels.91The White House, Building a Clean Energy Economy: A Guidebook to the Inflation Reduction Act’s Investments in Clean Energy and Climate Action 5–6 (2023); Summary: The Inflation Reduction Act of 2022, Senate Democrats, https://www.democrats.senate.gov/imo/media/doc/inflation_reduction_act_one_page_summary.pdf [https://perma.cc/Z4ED-W32A]. The Act reinforces the EPA’s authority to regulate GHGs under the Clean Air Act, amending sections of the Clean Air Act to define “greenhouse gas” to include “the air pollutants carbon dioxide, hydrofluorocarbons, methane, nitrous oxide, perfluorocarbons, and sulfur hexafluoride.”92Inflation Reduction Act of 2022, Pub. L. No. 117-169, § 132(d)(4), 136 Stat. 1818, 2067. It also grants money under the Clean Air Act for any project that “reduces or avoids greenhouse gas emissions and other forms of air pollution.”93Id. § 134(c)(3)(A), 136 Stat. 1818, 2064. This language supports that Congress fully intends to include GHGs in the Clean Air Act and that California is acting within the scope of the Clean Air Act in implementing its forward-looking motor vehicle emissions regulations.

C.  Pending Lawsuit—Ohio v. EPA

Similar to its prior motor vehicle regulations, California will need to request a preemption waiver from the EPA under section 209(b)(1) of the Clean Air Act to regulate post-2025 vehicles. In the meantime, the Biden administration’s EPA’s latest March 2022 waiver decision prompted Republican-led states and private petitioners to challenge the constitutionality of the Clean Air Act waiver provision, making the case highly relevant for California’s ability to regulate motor vehicle emissions in the future.94Brief for Petitioners, supra note 19, at 28. In May 2022, seventeen states filed a lawsuit in the U.S. Court of Appeals for the D.C. Circuit (Ohio v. EPA), claiming, inter alia, that the section 209(b)(1) waiver provision violates the equal sovereignty principle because it limits state political authority unequally by allowing only California to set new car emission standards and “exercise sovereign authority that section 209(a) takes from every other State.”95Id. Under this principle, the petitioners alleged, Congress cannot give only some states favorable treatment of sovereignty authority, as it has done with California.96Id. at 26. Even if section 209(b)(1) allowed California to regulate unique state-specific issues, the petitioners argued that the waiver would still be unconstitutional because it allows California to regulate GHGs to address climate change, which is not a problem unique to California.97Id. at 13. The petitioners disagreed with the Biden administration’s EPA’s statement that “California is particularly impacted by climate change,”982022 Waiver Reconsideration, 87 Fed. Reg. at 14363. arguing that other states will be impacted just as much, if not more, from climate change.99Brief for Petitioners, supra note 19, at 32.

The petitioner states also took issue with the idea of giving one state power to regulate a major national industry.100“A federal law giving one State special power to regulate a major national industry contradicts the notion of a Union of sovereign States.” Id. at 29–30. The states argued that California’s “special treatment” under the Clean Air Act—giving California special power to regulate a major national industry and exercise sovereign authority that the Act withdraws from every other state, when California has no unique interest101Id. at 26.—violates the Constitution’s intent to hold all states equal.102Id. at 30. “Instead of allowing all States with a unique environmental concern to seek a waiver, it accords special treatment to a category of States defined to forever include only California and to forever exclude all other States, without regard to whether other States face their own localized environmental concerns.” Id. at 30. In a separate brief, a group of private petitioners, including the American Fuel & Petrochemical Manufacturers and Clean Fuels Development Coalition, argued that the equal sovereignty principle does not allow the federal government to give only one state the authority to regulate national and international issues.103Initial Brief for Private Petitioners at 15, Ohio v. EPA, No. 22-1081 (D.C. Cir. Oct. 24, 2022). They claimed that any mandate to shift the nation’s automobile fleet to electric vehicles must come from Congress, because such a shift would “substantially restructure the American automobile market, petroleum industry, agricultural sectors, and the electric grid, at enormous cost and risk.”104Id. at 23. The private petitioners cited the recent West Virginia v. EPA decision, which essentially restricted the EPA’s authority to regulate GHG emissions from power plants.105See id. at 23; West Virginia v. EPA, 142 S. Ct. 2587, 2612, 2615–16 (2022). Applying the major questions doctrine,106The major questions doctrine states that if an agency seeks to decide an issue of major national significance—that is, in cases where the “history and breadth of the authority” an agency asserts or the “economic and political significance” of that assertion is extraordinary—its action must be supported by clear congressional authorization. Id. at 2607–08. See Kate R. Bowers, Cong. Rsch. Serv., IF12077, The Major Questions Doctrine 1 (2022) (providing an overview of the major questions doctrine). the Court held that the EPA must point to “clear congressional authorization”107 West Virginia, 142 S. Ct. at 2609 (quoting Util. Air Regul. Grp. v. EPA, 573 U.S. 302, 324 (2014)). to justify its regulatory authority in “extraordinary cases” when the EPA asserts broad authority in an area of “economic and political significance.”108West Virginia, 142 S. Ct. at 2608–09 (quoting FDA v. Brown & Williamson Tobacco Corp., 529 U.S. 120, 159–60 (2000)). The case centers around the Clean Power Plan, a regulation the EPA issued in 2015 that would have curbed carbon emissions from existing coal and gas plants via “‘generation shifting from higher-emitting to lower-emitting’ producers of electricity.” Id. at 2603 (quoting Carbon Pollution Emission Guidelines for Existing Stationary Sources: Electric Utility Generating Units, 80 Fed. Reg. 64728 (Oct. 23, 2015) (to be codified at 40 C.F.R. pt. 60)). The decision was the first time the Supreme Court has used the term “major questions doctrine” in a majority opinion. Bowers, supra note 106, at 2. The Court concluded that the EPA does not have the authority to “substantially restructure the American energy market . . . .”109West Virginia, 142 S. Ct. at 2610. If the EPA cannot upend energy generation in the country, as West Virginia v. EPA held, then, the petitioners argued, California similarly cannot “upend the transportation and energy sectors.”110Initial Brief for Private Petitioners, supra note 104, at 19–20. The petitioners further argued that section 177 also allows California to shape national industries, which may burden the states that decline to adopt California’s standards.111Id. at 54.

On the other hand, several electric utility providers, clean energy industry groups, and auto manufacturers have backed California.112Goldberg, supra note 13. A few automakers have indicated that they support the more stringent California standards. In July 2019, CARB reached a voluntary agreement with four major automakers—BMW of North America, Ford, Honda, and Volkswagen Group of America—to adopt a modified version of the GHG standards.113California and Major Automakers Reach Groundbreaking Framework Agreement on Clean Emission Standards, Cal. Air Res. Bd. (July 5, 2019), https://ww2.arb.ca.gov/news/california-and-major-automakers-reach-groundbreaking-framework-agreement-clean-emission [https://perma.cc/52VH-PCLS]. Building on this voluntary framework, in 2020, Volvo joined the four automakers in agreeing to a 17% emissions cut through the 2026 model year.114Framework Agreements on Clean Cars, Cal. Air Res. Bd. (Aug. 17, 2020), https://ww2.arb.ca.gov/news/framework-agreements-clean-cars [https://perma.cc/EN78-JR87]. The automakers filed a motion to intervene to defend the EPA’s March 2022 decision.115Ford Motor Co., Volkswagen Grp. of Am., Inc., BMW of N. Am., LLC, Am. Honda Motor Co., Inc., and Volvo Car USA LLC, Motion to Intervene in Support of Respondents, Ohio v. EPA, No. 22-1081 (D.C. Cir. June 7, 2022).

To date, the Supreme Court has not addressed the constitutionality of the Clean Air Act under the equal sovereignty principle. In its 2019 decision revoking the 2013 California waiver, the Trump administration’s EPA interpreted the statutory criteria in the context of the equal sovereignty principle, explaining that section 209(b)(1) provides “extraordinary treatment” to California and therefore should be interpreted to require a “state-specific particularized” pollution problem.1162019 Waiver Withdrawal, 84 Fed. Reg. at 51340. In contrast, in its 2022 waiver grant, the Biden administration’s EPA noted that it has historically declined to consider constitutional issues, reviewing the waiver solely based on the section 209(b)(1) criteria because the statute and legislative history reflect a broad policy of deference to California to address its air quality problems.1172022 Waiver Reconsideration, 87 Fed. Reg. at 14376. This interpretation has been upheld by the U.S. Court of Appeals for the D.C. Circuit. See Motor & Equip. Mfrs. Ass’n v. EPA (MEMA I), 627 F.2d 1095, 1115 (D.C. Cir. 1979) (declining to consider whether California standards are constitutional); Am. Trucking Ass’ns. v. EPA, 600 F.3d 624, 628 n.1 (D.C. Cir. 2010) (declining to express a view on a constitutional challenge to the California standards). In both cases, the Court upheld prior EPA decisions to not consider constitutional objections. Although equal sovereignty presented a new constitutional argument, the EPA limited its role in evaluating waiver requests to “the criteria that Congress directed EPA to review.”1182022 Waiver Reconsideration, 87 Fed. Reg. at 14376. Nevertheless, the Biden administration’s EPA briefly addressed the equal sovereignty principle, arguing that the waiver does not impose a burden on any state and that Section 177, in enabling other states to adopt California’s standards, undermines the notion that the section 209(b)(1) waiver treats California in an extraordinary manner.119Id. at 14356. Rather, in deliberately compromising between having one national standard and fifty different state standards by authorizing just two—the federal standard and California’s standards—Congress allowed California to be a “laboratory for innovation” and address the state’s extraordinary pollution problems, while ensuring that automakers were not overburdened with varying state standards.120Id. at 14360, 14377.

D.  California’s Advanced Clean Cars II Regulations

California recently promulgated the Advanced Clean Cars II (“ACC II”) regulations in the shadow of the pending Ohio v. EPA lawsuit. ACC II stems from an executive order Governor Gavin Newsom signed in September 2020 directing CARB to develop regulations contributing to the goal that 100% of in-state sales of new passenger cars and trucks will be zero-emission by 2035.121Cal. Exec. Order No. N-79-20 (Sept. 23, 2020), https://www.gov.ca.gov/wp-content/uploads/2020/09/9.23.20-EO-N-79-20-Climate.pdf [https://perma.cc/F4SE-B5AB]. As a point of comparison, in 2022, nearly 19% of all new light-duty vehicles sold in the state were electric vehicles. New ZEV Sales in California, Cal. Energy Comm’n, https://www.energy.ca.gov/data-reports/energy-almanac/zero-emission-vehicle-and-infrastructure-statistics/new-zev-sales [https://perma.cc/TDY9-TXST]. As a result of the executive order, on August 25, 2022, CARB promulgated a new regulation, the ACC II program, phasing out all sales of new fossil fuel cars by 2035.122Cal. Air Res. Bd., supra note 1. The regulation requires that automakers increase the percentage of electric vehicles progressively, nearly tripling it to 35% by 2026 and reaching 100% by 2035 (see Figure 1).123Cal. Air Res. Bd., supra note 3.

Figure 1.  Percentage of new vehicle sales that must be zero-emission vehicles

The ACC II regulations amend the ZEV and LEV standards for model years 2026–2035,124Cal. Air Res. Bd., supra note 77. The ACC II regulations: (1) amend the ZEV regulation to require an increasing number of zero-emission vehicles, and rely on advanced vehicle technologies, including battery-electric, hydrogen fuel cell electric and plug-in hybrid electric vehicles, to meet air quality and climate change emissions standards; and (2) amend the LEV regulations to include increasingly stringent standards for gasoline cars and heavier passenger trucks to continue to reduce smog-forming emissions while the sector transitions toward 100% electrification by 2035. Cal. Air Res. Bd., supra note 1. following the ACC I regulations, which address model year 2015–2025 vehicles.125Cal. Air Res. Bd., supra note 1. CARB estimates that the new regulations will reduce vehicle GHG emissions by more than 50% by 2040.126Goldberg, supra note 13. Thus, the decision from Ohio v. EPA will have implications for California’s ability to implement standards including the ACC II program going forward.

E.  The Equal Sovereignty Principle

The Supreme Court didn’t develop the equal sovereignty principle as a meaningful concept until Shelby County v. Holder in 2013,127Shelby County v. Holder, 570 U.S. 529, 540 (2013); see Equal Sovereignty Five Years After Shelby County, Harv. C.R.-C.L. L. Rev.: Amicus Blog (Oct. 31, 2018), https://harvardcrcl.org/equal-sovereignty-five-years-after-shelby-county [https://perma.cc/S5G8-QSAQ]. in which the Supreme Court held a statute (the Voting Rights Act) unconstitutional based on the equal sovereignty principle for the first time. The Court did not clarify what constitutional provision this principle is based on.128See Amdt 10.4.3 Equal Sovereignty Doctrine, Const. Annotated, https://constitution.congress.gov/browse/essay/amdt10-4-3/ALDE_00013628 [https://perma.cc/US7J-4YU9]. Although the Constitution requires equal treatment among the states in particular contexts,129See, e.g., U.S. Const. art. I, § 3, cl. 1 (“The Senate of the United States shall be composed of two Senators from each State . . . .”); U.S. Const. art. I, § 8, cl. 1 (requiring “Duties, Imposts and Excises” to be “uniform throughout the United States”); U.S. Const. art. I, § 8, cl. 4 (requiring “a uniform Rule of Naturalization” and “uniform Laws on the subject of Bankruptcies throughout the United States”); U.S. Const. art. I, § 9, cl. 6 (“No Preference shall be given by any Regulation of Commerce or Revenue to the Ports of one State over those of another . . . .”); U.S. Const. art. IV, § 1 (Full Faith and Credit Clause – “Full Faith and Credit shall be given in each State to the public Acts, Records, and judicial Proceedings of every other State.”); U.S. Const. art. IV, § 2, cl. 1 (Privileges and Immunities Clause – “The Citizens of each State shall be entitled to all Privileges and Immunities of Citizens in the several States.”); U.S. Const. art. V (“[N]o State, without its Consent, shall be deprived of its equal Suffrage in the Senate.”); U.S. Const. amend. XI. no provision explicitly requires Congress to treat all states equally as a general matter.130See Leah M. Litman, Inventing Equal Sovereignty, 114 Mich. L. Rev. 1207, 1230–32 (2016); Thomas B. Colby, In Defense of the Equal Sovereignty Principle, 65 Duke L.J. 1087, 1099–1100 (2016). This absence of an explicit statement could mean that the founders did not intend to establish a generally applicable equal sovereignty principle.131See Final Brief for Respondents at 33, Ohio v. EPA, No. 22-1081 (D.C. Cir. Mar. 20, 2023). Critics of Shelby County have claimed that the Supreme Court invented the equal sovereignty principle to achieve political ends.132See Abigail B. Molitor, Understanding Equal Sovereignty, 81 U. Chi. L. Rev. 1839, 1840 (2014); Litman, supra note 130. Judge Richard Posner, Chief Judge of the Seventh Circuit Court of Appeals, wrote regarding the equal sovereignty principle: “This is a principle of constitutional law of which I had never heard—for the excellent reason that . . . there is no such principle . . . . The opinion [Shelby County] rests on air.” Richard A. Posner, The Supreme Court and the Voting Rights Act: Striking Down the Law Is All About Conservatives’ Imagination, Slate (June 26, 2013, 12:16 AM), https://slate.com/news-and-politics/2013/06/the-supreme-court-and-the-voting-rights-act-striking-down-the-law-is-all-about-conservatives-imagination.html [https://perma.cc/P7WJ-62A7]. Other scholars argue that questions about the sovereign power of the states have existed since the drafting of the U.S. Constitution.133See Molitor, supra note 132, at 1877; Colby, supra note 130, at 1102; Valerie J.M. Brader, Congress’ Pet: Why the Clean Air Act’s Favoritism of California Is Unconstitutional Under the Equal Footing Doctrine, 13 Hastings W.-Nw. J. Env’t L. & Pol’y 119, 151 (2007); Jeffrey M. Schmitt, In Defense of Shelby County’s Principle of Equal State Sovereignty, 68 Okla. L. Rev. 209, 238 (2016). Many scholars agree there is some support for the principle in the historical record and constitutional doctrine, but they doubt that is sufficient for it to be considered a “fundamental” principle, as Shelby County claims.134See Molitor, supra note 132, at 1841; Litman, supra note 130, at 1212; David Kow, An “Equal Sovereignty” Principle Born in Northwest Austin, Texas, Raised in Shelby County, Alabama, 16 Berkeley J. Afr.-Am. L. & Pol’y 346, 375 (2015). This Section traces the history of how courts have applied the equal sovereignty principle, from the context of admitting new states into the Union to voting rights.

1.  Origins: The Equal Footing Doctrine—New Admission of States

The equal sovereignty principle dates back to the equal footing doctrine referenced in Article IV, Section 3 of the Constitution: “New States may be admitted by the Congress into this Union; but no new State shall be formed or erected within the Jurisdiction of any other State . . . without the Consent of the Legislatures of the States concerned as well as of the Congress.”135U.S. Const. art. IV, § 3. The Northwest Ordinance of 1787, which provided a path toward statehood for the territories northwest of the Ohio River,136These territories would later become Illinois, Indiana, Michigan, Ohio, Wisconsin, and part of Minnesota. The Northwest Ordinance of 1787, U.S. H.R.: Hist., Art & Archives, https://history.house.gov/Historical-Highlights/1700s/Northwest-Ordinance-1787/ [https://perma.cc/CLG2-V2ZA]. further required that these states be admitted “on an equal footing with the original States in all respects whatever,” on the condition that the new state constitutions and governments were “republican, and in conformity to the principles contained in these articles . . . .”137Ordinance for the Government of the Territory of the United States North-West of the River Ohio art. V (1787), https://www.archives.gov/milestone-documents/northwest-ordinance [https://perma.cc/2ZUF-U5DT]. The act also banned slavery in the new territories but allowed for the return of fugitive slaves. Id., art. VI. Professor Litman argues, however, that the Northwest Ordinance’s meaning is unclear because “equal footing” did not necessarily promise new states the same legislative sovereignty as the original states, but rather just that new states would receive fair representation in Congress. Litman, supra note 130, at 1235–36. Additionally, Litman notes that the Northwest Ordinance actually broadened Congress’s powers over the would-be states, resulting in different treatment of those states, since it prohibited religious discrimination and slavery in the new states. Id. James Madison inferred that Congress would determine whether newly admitted states have the same “legislative sovereignty” as the original states. Id.

Several court cases also interpret the Constitution to support the equal sovereignty principle. Pollard’s Lessee v. Hagan held that Congress must admit every state into the Union on the same terms and with the same powers as the original states.138“The new states have the same rights, sovereignty, and jurisdiction [over the shores of navigable waters] as the original states.” Pollard’s Lessee v. Hagan, 44 U.S. 212, 230 (1845). Every state must be “admitted into the union on an equal footing with the original states,139Id. at 216. with “equal sovereign rights.”140Id. at 231. Further, the court held that “no compact” can “diminish or enlarge” the rights a state has when it enters the Union.141Id. at 229. Northwest Austin v. Holder referenced this case as support for the historic tradition that all states enjoy equal sovereignty.142Nw. Austin Mun. Util. Dist. No. 1 v. Holder, 557 U.S. 193, 203 (2009) (citing United States v. Louisiana, 363 U.S. 1, 16 (1960) (citing Pollard’s Lessee v. Hagan, 44 U.S. 212, 223 (1845))). Coyle v. Smith held that states, not Congress, have sovereignty to choose where to locate their state capital: the United States “was and is a union of States, equal in power, dignity and authority, each competent to exert that residuum of sovereignty not delegated to the United States by the Constitution itself.”143Coyle v. Smith, 221 U.S. 559, 567 (1911). No state is “less or greater . . . in dignity or power” than another.144Id. at 566. Thus, Congress may not unequally limit or expand the states’ political and sovereign power.145See Stearns v. Minnesota, 179 U.S. 223, 245 (1900) (“It has often been said that a State admitted into the Union enters therein in full equality with all the others, and such equality may forbid any agreement or compact limiting or qualifying political rights and obligations . . . .”). Indeed, “the constitutional equality of the States is essential to the harmonious operation of the scheme upon which the Republic was organized.”146Coyle, 221 U.S. at 580. Thus, these cases establish the origins of the equal sovereignty principle in the admission of new states into the Union.

2.  Equal Sovereignty Applied to Voting Rights

When the equal sovereignty principle was brought up in the context of the Voting Rights Act, courts had to determine whether the principle applied outside the state admission context.

Congress designed the Voting Rights Act of 1965 to address continuing voting discrimination after the Civil War.147South Carolina v. Katzenbach, 383 U.S. 301, 308 (1966). The Fifteenth Amendment to the Constitution, ratified in 1870, prohibited voting discrimination based on race,148See id. at 310; U.S. Const. amend. XV, § 1. and Congress subsequently enacted the Enforcement Act of 1870, which prohibited obstruction of the exercise of the right to vote.149See Katzenbach, 383 U.S. at 310; Enforcement Act of 1870, ch. 114, 41st Congress, Sess. II. However, enforcement of the law was ineffective, and throughout Reconstruction, many southern states continued to enact tests designed to prevent Black people from voting.150See Katzenbach, 383 U.S. at 310–11. Literacy tests disproportionately affected African Americans due to the high illiteracy rates in comparison with Whites. At the same time, grandfather clauses, property qualifications, character tests, and interpretation requirements were employed to “assure that white illiterates would not be deprived of the franchise.” Id. at 311. To address this continuing discrimination, section 5 of the Voting Rights Act established a preclearance requirement, mandating that the federal government approve all new voting regulations to ensure that they did not perpetuate racial discrimination.151Voting Rights Act of 1965, Pub. L. No. 89-110, § 5, 79 Stat. 437, 439. However, the preclearance requirement only applied to states with a history of voting discrimination, as determined by the coverage formula in section 4 of the Voting Rights Act.152The coverage formula established that if the state used a law like a literacy or character test to keep people from registering to vote as of November 1, 1964, and less than 50% of the eligible voting population was registered to vote on November 1, 1964 or voted in the presidential election of November 1964, then the state was subject to preclearance. Voting Rights Act of 1965, Pub. L. No. 89-110, § 4(b), 79 Stat. 437, 438. The coverage formula implicated states located primarily in the South; thus, a select group of states were subject to more stringent requirements than other states when seeking to change their voting laws.

In its 1966 decision in South Carolina v. Katzenbach, the Supreme Court rejected the notion that the equal sovereignty principle prohibited differential treatment in the voting rights context. The Court held that the equal sovereignty principle only applied to situations involving the admission of new states, not the Voting Rights Act: “The doctrine of the equality of States . . . applies only to the terms upon which States are admitted to the Union, and not to the remedies for local evils which have subsequently appeared.”153Katzenbach, 383 U.S. at 328–29. The Court observed that Congress passed the Voting Rights Act in response to the “insidious and pervasive evil” of racial discrimination in voting,154Id. at 309. and thus held that the Voting Rights Act was a constitutional and appropriate means for carrying out the Fifteenth Amendment.155Id. at 328–29.

Fourteen years later in City of Rome v. United States, the Supreme Court again upheld the Voting Rights Act as constitutional, finding that the Reconstruction Amendments were “specifically designed as an expansion of federal power and an intrusion on state sovereignty,” and thus, Congress had the authority to regulate state and local voting.156City of Rome v. United States, 446 U.S. 156, 179 (1980). The Court cited Fitzpatrick v. Bitzer, which held that the principle of state sovereignty embodied by the Eleventh Amendment is “necessarily limited by the enforcement provisions of section 5 of the Fourteenth Amendment.”157Id. at 156–58 (citing Fitzpatrick v. Bitzer, 427 U.S. 445, 456 (1976)). However, the Court would later apply the equal sovereignty principle to invalidate part of the Voting Rights Act.

3.  Shelby County v. Holder—Equal Sovereignty as a General Principle

Only two Supreme Court cases discuss equal sovereignty as a general principle.158Molitor, supra note 132, at 1879. Northwest Austin v. Holder,159Nw. Austin Mun. Util. Dist. No. 1 v. Holder, 557 U.S. 193, 203 (2009). though still a voting rights case, applied the equal sovereignty principle more broadly in 2009, laying the foundation for Shelby County v. Holder160Shelby County v. Holder, 570 U.S. 529, 544 (2013). to overrule Voting Rights Act section 4 in 2013.161See Molitor, supra note 132, at 1878 (“Since Shelby County, only one court has issued an opinion dealing with equal sovereignty [NCAA v. New Jersey, a Third Circuit case].”).

In Northwest Austin, the Supreme Court observed that the section 4 coverage formula of the Voting Rights Act went against the “historic tradition that all the States enjoy ‘equal sovereignty’ ” by differentiating between the states.162Nw. Austin, 557 U.S. at 203 (citing United States v. Louisiana, 363 U.S. 1, 16 (1960) (citing Lessee of Pollard v. Hagan, 3 How. 212, 223 (1845))). The Court acknowledged that differentiating between states is sometimes justified, citing Katzenbach as an example.163Id.  (citing South Carolina v. Katzenbach, 383 U.S. 301, 328–29 (1966)). However, it held that departing from “the fundamental principle of equal sovereignty requires a showing that a statute’s disparate geographic coverage is sufficiently related to the problem that it targets.”164Id. Thus, the equal sovereignty principle limits Congress’s ability to subject different states to unequal burdens, at least without sufficient justification.165Amdt 10.4.3 Equal Sovereignty Doctrine, Const. Annotated, https://constitution.congress.gov/browse/essay/amdt10-4-3/ALDE_00013628 [https://perma.cc/US7J-4YU9]. The Court also noted that the Act “imposes current burdens and must be justified by current needs.”166Nw. Austin, 557 U.S. at 203. While the Court ultimately resolved the case on statutory grounds,167Id. at 206–11. it expressed concern that sections 4 and 5 of the Voting Rights Act raised “serious constitutional questions.”168Id. at 204. The Court observed that improved conditions in the South since 1965 may distinguish the case from Katzenbach because current conditions in 2009 may no longer reflect the discriminatory state actions that Congress meant for section 5 to address, and cited a lower racial gap in voter registration as an example to show that the coverage formula may rely on outdated statistics.169Id. at 202–04 (2009). The Court notes that “[v]oter turnout and registration rates now approach parity[,]” “[b]latantly discriminatory evasions of federal decrees are rare,” and “minority candidates hold office at unprecedented levels.” Id. at 202. The Court also observed that the Voting Rights Act’s preclearance requirements “authorize[d] federal intrusion into sensitive areas of state and local policymaking” and imposed “substantial ‘federalism costs.’ ”170Id. at 202.

These concerns formed the basis for Shelby County to hold that section 4 of the Voting Rights Act was unconstitutional because it departed from the “fundamental principle” of equal sovereignty.171Shelby County v. Holder, 570 U.S. 529, 544 (2013). The Supreme Court found the “fundamental principle of equal sovereignty” to be “highly pertinent in assessing subsequent disparate treatment of States.”172Id. The Court adopted the guidelines Northwest Austin set—namely, that the Voting Rights Act “imposes current burdens and must be justified by current needs,” and that “a departure from the fundamental principle of equal sovereignty requires a showing that a statute’s disparate geographic coverage is sufficiently related to the problem that it targets.”173Id. at 542; see Nw. Austin, 557 U.S. at 203. The Court also distinguished the case from Katzenbach. Whereas in Katzenbach, the coverage formula was “relevant to the problem” of voting discrimination at the time,174Shelby County, 570 U.S. at 551–52; see South Carolina v. Katzenbach, 383 U.S. 301, 301 (1966). here, the coverage formula was not updated to reflect contemporary improvements in voting participation, including higher voter registration and turnout numbers.175Shelby County, 570 U.S. at 547–49, 551. The Court concluded that Congress did not sufficiently justify its reauthorization of the “extraordinary and unprecedented features” of the Voting Rights Act;176Id. at 549. thus, the Court held that the coverage formula no longer met the test introduced in Northwest Austin.177Id. at 551.

Shelby County, the only Supreme Court case to apply the test established in Northwest Austin, gave little guidance on how to apply the equal sovereignty principle in future cases, other than indicating that the law should rely on “current data reflecting current needs” when the degree of voting discrimination that prompted the original passage of the Voting Rights Act had changed.178Id. at 552–53. The Supreme Court has not decided an equal sovereignty challenge since Shelby County, leaving lower courts to interpret how to apply the equal sovereignty principle outside the voting rights context.

II.  APPLYING THE SHELBY COUNTY TEST TO THE CLEAN AIR ACT

Under the Northwest Austin test that Shelby County applied (the “Shelby County test”), the statute “must be justified by current needs,” and if federal legislation departs from the “fundamental principle of equal sovereignty,” it “requires a showing that a statute’s disparate geographic coverage is sufficiently related to the problem that it targets.”179Id. at 542; see Nw. Austin, 557 U.S. 193, 203 (2009). This Part argues that the equal sovereignty principle likely does not apply to the Clean Air Act, thus the Shelby County test should not even apply. But even if it were to apply and the Shelby County test is triggered, this Part concludes that the principle does not invalidate section 209(b)(1) of the Clean Air Act because California’s current needs continue to justify its differential treatment. California’s unique exemption is sufficiently related to the public health problem that the Clean Air Act waiver provision targets; allowing California broad discretion to regulate motor vehicle emissions directly contributes to Congress’s goal of addressing public health threats from motor vehicle pollution in the state.

A.  The Equal Sovereignty Principle Likely Does Not Apply to the Clean Air Act

This Section argues that the scope of the Shelby County test is limited and likely does not apply to the Clean Air Act. Shelby County emphasizes that the equal sovereignty principle applies to federal laws that “authorize[] federal intrusion into sensitive areas of state and local policymaking.”180Shelby County, 570 U.S. at 545 (citing Lopez v. Monterey County, 525 U.S. 266, 282 (1999)). The Supreme Court thus applied the equal sovereignty principle to the Voting Rights Act because it determined that election regulation was a sensitive area of state policymaking. Highlighting the “extraordinary” nature of the Voting Rights Act’s preclearance provisions,181Id. the Court noted that the law suspends “all changes to state election law—however innocuous—until they have been precleared by federal authorities . . . .”182Id. at 544 (citing Nw. Austin, 557 U.S. at 202). The federal government must explicitly grant states permission to implement voting laws that they “would otherwise have the right to enact and execute on their own . . . .”183Id. Because the Voting Rights Act intruded into a sensitive area of state policymaking that had traditionally been the exclusive province of the states, the Court limited Congress’s authority under the Fifteenth Amendment to restrict states’ election procedures disparately.

Professor Leah Litman goes even farther to posit that only federal action that lessens the dignity of a state or group of states triggers the Shelby County conception of equal sovereignty.184Litman, supra note 130, at 1214. Under this narrower interpretation, Litman argues that laws will violate equal sovereignty only if they single out particular states that have behaved in morally-blameworthy ways, limiting the scope of the principle to legislation enacted under the Reconstruction Amendments.185Id. at 1214–15. Under this interpretation, the equal sovereignty principle primarily serves as a check on the Fourteenth and Fifteenth Amendments and should only apply in cases similar to those involving voting rights, in which the dignity of human beings is at stake.186Id.

Since Shelby County, a few weak equal sovereignty claims have been made in the lower courts in areas outside of voting rights, and the courts have distinguished these cases from Shelby County. For example, in Mayhew v. Burwell, the U.S. Court of Appeals for the First Circuit held that the equal sovereignty principle does not apply to Medicaid laws.187In Mayhew v. Burwell, the U.S. Court of Appeals for the First Circuit held that the Affordable Care Act (“ACA”) did not violate equal sovereignty even though it prevented Maine from “design[ing] its [own] Medicaid laws in ways that many of its sister States remain[ed] free to do.” Mayhew v. Burwell, 772 F.3d 80, 93 (1st Cir. 2014). The court reasoned that the ACA did not intrude into an area of authority traditionally occupied by the states because it governed Maine’s administration of a federal program that is primarily funded by the federal government. Id. at 95. Thus, the statute at issue “does not similarly effect a federal intrusion into a sensitive area of state or local policymaking.” Id. at 93. Perhaps most relevant to the Clean Air Act waiver is National Collegiate Athletic Association (NCAA) v. Governor of New Jersey, which addressed the constitutionality of the Professional and Amateur Sports Protection Act of 1992 (“PASPA”).188Nat’l Collegiate Athletic Ass’n v. Governor of N.J., 730 F.3d 208, 214 (3d Cir. 2013). PASPA prohibits states from licensing sports gambling, except for states that had gambling operations prior to the Act’s passage, which only includes Nevada.189Id. at 214–15; see 28 U.S.C. § 3702, 3704. The U.S. Court of Appeals for the Third Circuit determined that the equal sovereignty principle does not apply to PASPA, distinguishing the Voting Rights Act from PASPA by finding that regulating gambling via the Commerce Clause is “not of the same nature” as regulating elections via the Reconstruction Amendments.190Nat’l Collegiate Athletic Ass’n, 730 F.3d at 238. The court held that the Commerce Clause allowed Congress to enact laws “aimed at matters of national concern and finding national solutions will necessarily affect states differently,”191Id. such that federal Commerce Clause regulation “does not require geographic uniformity.”192Id. (citing Morgan v. Virginia, 328 U.S. 373, 388 (1946)). The court found that applying Shelby County to all situations is “overly broad” and that the equal sovereignty principle does not apply outside “the context of ‘sensitive areas of state and local policymaking.’ ”193Id. at 238–39 (citing Shelby County v. Holder, 570 U.S. 529, 545 (2013)).

Similar to PASPA, Congress acted pursuant to its Commerce Clause authority in passing the Clean Air Act to regulate motor vehicle emissions; thus, Congress is exercising the federal power of regulating interstate commerce and can treat states differently in the process.194See Vikram David Amar, Why the Clean Air Act’s Special Treatment of California is Permissible Even in Light of the Equal-Sovereignty Notion Invoked in Shelby County, Justia: Verdict (Aug. 2, 2022), https://verdict.justia.com/2022/08/02/why-the-clean-air-acts-special-treatment-of-california-is-permissible-even-in-light-of-the-equal-sovereignty-notion-invoked-in-shelby-county [https://perma.cc/EYD8-H26N] (“[T]he Clean Air Act was enacted under Congress’s Commerce Clause powers, a provision that decidedly does not require geographic uniformity”); Final Brief for Respondents, supra note 131, at 32–35. The Clean Air Act likely does not intrude into “sensitive areas of state and local policymaking” as the Voting Rights Act does. Regulating motor vehicles has not traditionally been the exclusive province of the states. Three agencies set federal and state vehicle emissions standards: the EPA, the National Highway Traffic Safety Administration, and CARB.195Federal Vehicle Standards, Ctr. for Climate & Energy Sols., https://www.c2es.org/content/regulating-transportation-sector-carbon-emissions [https://perma.cc/BK55-6TCA]. Section 209(a) of the Clean Air Act explicitly provides for federal preemption, prohibiting states from adopting their own motor vehicle regulations.19642 U.S.C. § 7543(a). Regulating motor vehicle emissions affects interstate commerce because air pollution crosses state borders.197S. Allan Adelman, Control of Motor Vehicle Emissions: State or Federal Responsibility? 20 Cath. U. L. Rev. 157, 158, 163–64 (1970). Thus, like PASPA, the Clean Air Act does not intrude into a sensitive area of policymaking traditionally occupied by the states.

At its core, the outcome the petitioners demand in Ohio v. EPA is inconsistent with the fundamental principle of equal sovereignty. Without the waiver, the Clean Air Act defaults to only federal standards and federal preemption, leaving states with no choice but to adopt the federal standard. Thus, invalidating the California waiver—as petitioners seek to do—gives states fewer choices. It fails to promote the principle of equal sovereignty, which arguably protects the power of the states to enact policies that differ from those of the federal government.198Schmitt, supra note 133, at 262; see infra Section I.E.1; 2022 Waiver Reconsideration, supra note 86, at 14360 (“Indeed, if section 209(b) is interpreted to limit the types of air pollution that California may regulate, it would diminish the sovereignty of California and the states that adopt California’s standards pursuant to section 177 without enhancing any other state’s sovereignty.”). In her amicus brief, Professor Litman noted that the petitioners’ invocation of the equal sovereignty principle is inconsistent with its history because the petitioners’ arguments would result in less authority and flexibility for the states, and more coercive authority for the federal government.199Brief for Professor Leah M. Litman as Amici Curiae Supporting Respondents at 2, Ohio v. EPA, No. 22-1081 (D.C. Cir. Jan. 20, 2023). By allowing California to promulgate more stringent standards and allowing other states to choose between the federal and California standards, Congress has offered those states more options, not fewer. This is likely not an abuse of state sovereignty.200Id. at 28. By arguing for an expansion of federal preemption, thereby preempting more state legislative and policy goals, the petitioners seek a result that does not promote state sovereignty and instead runs contrary to the equal sovereignty principle’s historical use as a limit on congressional power.201Id. at 5, 30; see infra Section I.E.1.

Congressional debates regarding California’s special status indicate that Congress clearly considered the equal sovereignty problem and rejected it. In 1970, members of the House of Representatives expressed concern that all states should have the “same right that the State of California has in setting standards that they deem necessary for the health and safety of their people.”202See 91 Cong. Rec. H19232 (daily ed. Jun. 10, 1970) (statement of Rep. Leonard Farbstein, New York). Representatives of other states, including Pennsylvania and New York, argued that their air quality problems were worse than California’s, so they too should have the power to create state regulations exceeding federal standards.203Pennsylvania “has had more deaths due to air pollution than any other State in the Nation” and “is interested in increasing its standards.” Id. at 19231. “New York has a problem with fog and smog that is just as bad as that condition which exists in California.” Id. at 19232. Thus, proper application of the equal sovereignty principle would allow all states to promulgate their own motor vehicles emissions regulations. Congress was more concerned about other states not being able to promulgate their own motor vehicles emissions standards than about California having special privileges. In contrast, in Ohio v. EPA, the petitioner states attempt to prevent California from enacting more stringent policies that could benefit other states, thus flipping the use of the equal sovereignty principle to make it more difficult for states to enact their own policies.

The Supreme Court has suggested in Shelby County that the equal sovereignty principle does not extend to all areas of the law, and this Section concludes that the equal sovereignty principle does not apply to the Clean Air Act. However, even if it were to apply, the Clean Air Act waiver provision passes the Shelby County test and remains constitutional, as analyzed in the next Section.

B.  Even if the Equal Sovereignty Principle Applies to the Clean Air Act, It Does Not Invalidate Section 209(b)(1) of the Clean Air Act

Even if the equal sovereignty principle were to apply to the Clean Air Act, the Clean Air Act waiver provision remains constitutional. Applying the Shelby County test, the Clean Air Act waiver likely departs from the “fundamental principle of equal sovereignty” in creating a differential in its treatment of states’ political authority. As a result, the “statute’s disparate geographic coverage” must be “sufficiently related to the problem that it targets.” This Section concludes that this criterion is met; thus, the waiver provision remains constitutional. Congress had strong justifications for granting California an exemption that continue to remain relevant. First, the Clean Air Act targets not only smog in one region of California, but also the broader problem of public health from automobile emissions. Second, allowing California to implement more stringent motor vehicle regulations would directly help address this broader problem. California faces new and increasingly formidable threats from climate change, which have exacerbated the existing problems that initially compelled California’s motor vehicle regulations. Allowing California broad discretion to regulate GHG emissions is directly related to Congress’s goal of addressing the public health threats from motor vehicle pollution in California because the effects of GHG emissions and smog are interrelated and affect one another. This Section thus concludes that California’s current needs continue to justify Congress’s differential treatment of California—maintaining, and perhaps even strengthening, section 209(b)’s relevance in the twenty-first century.

1.  By Treating States’ Political Authority Differently, the Clean Air Act Waiver Likely Violates the Equal Sovereignty Principle

The equal sovereignty principle does not require the federal government to treat states equally in every scenario, but requires that all states have equal political authority.204Schmitt, supra note 133, at 220. Black’s Law Dictionary defines “sovereignty” as “[s]upreme dominion, authority, or rule”205Sovereignty, Black’s Law Dictionary (11th ed. 2019). and “state sovereignty” as “[t]he right of a state to self-government; the supreme authority exercised by each state.”206State sovereignty, Black’s Law Dictionary (11th ed. 2019). The Court in Shelby County explained that “[s]tates retain broad autonomy . . . in structuring their governments and pursuing legislative objectives,”207Shelby County v. Holder, 570 U.S. 529, 543 (2013). referencing the Tenth Amendment and federalism principles as crucial in preserving the “integrity, dignity, and residual sovereignty of the States.”208Id. at 530 (citing Bond v. United States, 564 U.S. 211, 221 (2011)). In United States v. Texas, the Supreme Court noted that the equal footing doctrine applies to political rights and sovereignty, but not economic issues.209United States v. Texas, 339 US 707, 716 (1950). The Court observed that the equal footing doctrine was not designed to eliminate diversity in economic aspects such as area, location, and geology, but rather to “create parity as respects political standing and sovereignty.”210Id. Thus, Congress violates the equal sovereignty principle when it limits the political power of a particular subset of states.211Schmitt, supra note 133, at 220.

Legislation that prohibits some states but not others from enacting laws about the same topic likely would violate the equal sovereignty principle. For example, the Voting Rights Act limits only southern states’ ability to regulate elections and PASPA permits only Nevada to legalize sports betting;212Colby, supra note 130, at 1155. PASPA “does not merely regulate private conduct; it curtails the regulatory and revenue-raising authority of the states. It precludes non-exempted states from legalizing sports gambling . . . . Nevada may derive enormous financial benefits from casino sports book betting, but other states may not.” Id. thus, these laws would in theory violate the principle. Similarly, the Clean Air Act treats California’s sovereign authority differently from the other states. By permitting only California to regulate motor vehicles and promulgate new motor vehicles emissions standards, while limiting other states to either adopt the California or federal standards, the Clean Air Act waiver arguably limits other states’ rights to govern themselves in the area of motor vehicles, as well as transportation and energy more broadly. Rather than allow all states with certain air quality conditions to set regulations, the Clean Air Act allowed the state that first adopted its own motor vehicle regulations to continue setting the standard for new regulations.213See Brader, supra note 133, at 155–56. “The one state that had chosen to regulate in particular ways was given a power denied to all the states that had chosen not to exercise their equal right to do so . . . . These provisions are not about an inequality of economics or geography—they are about sovereignty.” Id. Thus, if we were to apply the equal sovereignty principle to the Clean Air Act, the Clean Air Act likely departs from the equal sovereignty principle by exhibiting disparate treatment of the states’ political authority pertaining to motor vehicle regulations.

2.  Nevertheless, the Clean Air Act Waiver Provision Remains Constitutional Because Its Disparate Geographic Coverage Favoring California Is “Sufficiently Related to the Problem that It Targets”

Violating the equal sovereignty principle does not automatically invalidate a law as unconstitutional. However, it triggers heighted scrutiny, meaning that Congress must justify the disparate treatment of the states as unequal sovereigns214See Nw. Austin Mun. Util. Dist. No. 1 v. Holder, 557 U.S. 193, 203 (2009) (“Distinctions can be justified in some cases.”). by showing that the differential treatment is sufficiently related to the problem the law is addressing.215Colby, supra note 130, at 1155–56. If the statute departs from the “fundamental principle of equal sovereignty,” it “requires a showing that a statute’s disparate geographic coverage is sufficiently related to the problem that it targets.” Shelby County v. Holder, 570 U.S. 529, 542 (2013) (citing Nw. Austin, 557 U.S. at 203 (2009)). This higher standard “ensures that when Congress limits the sovereign power of some of the states in ways that do not apply to others, it has a good reason to do so.”216Schmitt, supra note 133, at 213.

In Shelby County, the Supreme Court concluded that the coverage formula, while perhaps justified in 1965, was no longer justified in 2006 when Congress reauthorized the Voting Rights Act.217Shelby County, 570 U.S. at 551. Because the coverage formula continued to distinguish states “based on ‘decades-old data and eradicated practices,’ ” including the past use of literacy tests that “have been banned nationwide for over 40 years” and on racial disparity in “voter registration and turnout in the 1960s and early 1970s” that no longer persisted, the Court held that the 2006 reauthorization statute’s disparate geographic coverage was not sufficiently related to the problem of twenty-first century racial discrimination in voting that it targeted, so “current needs” no longer justified it.218Id. at 551–53. Thus, the Court found circumstances in 2013 to be sufficiently changed to render the coverage formula unconstitutional.219Id. at 550–53, 556–57; Molitor, supra note 132, at 1849–50.

Applying this line of reasoning to the Clean Air Act, the petitioners in Ohio v. EPA claim that because California has transitioned to regulating GHG emissions, the waiver provision is no longer sufficiently related to the problem that it targets because California’s standards are targeting climate change, which is global, not state-specific, in nature: “[C]limate change is not an acute California problem.”220Brief for Petitioners, supra note 19, at 30–31. This Section counteracts this argument and asserts that the waiver provision continues to be sufficiently related to the problem that it targets, distinguishing California’s motor vehicle regulations from the voting regulations at issue in Shelby County. First, the Clean Air Act targets not only smog in one region of California, but also the broader problem of public health from automobile emissions. Second, allowing California to implement more stringent motor vehicle regulations would directly help address this broader problem. California faces new and increasingly formidable threats from climate change that have exacerbated the existing problems that initially compelled California’s motor vehicle regulations. The effects of GHG and smog pollution are directly interrelated and affect one another; thus, addressing GHG emissions is directly related to Congress’s goal of addressing the public health threats from motor vehicle pollution in California. This Section therefore concludes that California’s current needs continue to justify the state’s differential treatment.

i.  The Clean Air Act Targets the Broad Problem of Public Health Threats from Automobile Emissions

How courts frame the problem that Congress is targeting can shape their determination of whether a statute is constitutional. In NCAA v. Governor of New Jersey, the U.S. Court of Appeals for the Third Circuit held that even if the equal sovereignty principle were to apply to Commerce Clause legislation, PASPA passed the Shelby County test because its “true purpose” was to “stop the spread of state-sanctioned sports gambling,” rather than eliminate it altogether.221Nat’l Collegiate Athletic Ass’n v. Governor of N.J., 730 F.3d 208, 239 (3d Cir. 2013). Because PASPA was drafted in neutral terms, any state that already supported gambling could continue to do so, and Congress likely knew that Nevada was the only state that had existing gambling operations.222“It shall be unlawful . . . to sponsor, operate, advertise, promote, license, or authorize by law or compact . . . .” 28 U.S.C. § 3702. However, § 3702 shall not apply to a state that conducted a gambling scheme “at any time during the period beginning January 1, 1976, and ending August 31, 1990 . . . .” 28 U.S.C. § 3704. “Nevada alone began permitting widespread betting on sporting events in 1949 . . . .” Nat’l Collegiate Athletic Ass’n, 730 F.3d at 215. PASPA’s disparate geographic coverage was therefore justified: “Targeting only states where the practice did not exist is . . . precisely tailored to address the problem.”223Nat’l Collegiate Athletic Ass’n, 730 F.3d at 239. If the court had defined the problem PASPA was targeting as eliminating all sports gambling, Nevada’s exemption would be harder to justify, and the statute would likely be unconstitutional for not being sufficiently related to the problem. However, because the court defined the problem as halting the spread of sports gambling, the Third Circuit’s analysis was a stronger one.

In Ohio v. EPA, the petitioners argue that the problem Congress designed the Clean Air Act to target was a narrow, California-specific problem.224Brief for Petitioners, supra note 19, at 30–31. However, while smog may have been the impetus for the legislation,225See H.R. Rep. No. 90-728, at 50 (1967) (recognizing the “critical concern of California for air pollution control, which is prompted especially by the acute susceptibility of the Los Angeles basin to concentrations of smog”). Congress also intended a broader goal of enabling California to use its developing expertise in vehicle pollution to develop innovative regulatory programs and serve as a leader in automobile emissions regulations.226See Chanin, supra note 63, at 716–17. In 1967, Congress acknowledged California’s serious air quality problems as well as its role as a laboratory for emissions control technology for the country.227See H.R. Rep. No. 90-728, at 96 (1967). The Senate Report concluded that with California’s experience in control systems and design, the waiver provision will allow California to “continue to be the testing area” for more stringent standards, potentially strengthening federal standards and benefiting all states.228S. Rep. No. 90-403, at 33 (1967).

Multiple instances from the Congressional Record suggest that the broader problem Congress intended to target was the public health threats caused by motor vehicle pollution.229See H.R. Rep. No. 90-728, at 3–8, 96 (1967); S. Rep. No. 90-403, at 32–33 (1967). Congress could have amended the Clean Air Act in 1977 to restrict the waiver provision. Instead, it ratified and strengthened the waiver by giving California the flexibility to adopt a complete program of motor vehicle emission controls.230Motor & Equip. Mfrs. Ass’n v. EPA (MEMA I), 627 F.2d 1095, 1110 (D.C. Cir. 1979) (citing H.R. Rep. No. 95-294, at 301–02 (1977); see infra Section I.A. The original 1967 waiver provision required the EPA Administrator to grant a waiver “unless he finds that such State does not require standards more stringent than applicable Federal standards . . . .”231Clean Air Act of 1967, Pub. L. No. 90-148, § 208(b), 81 Stat. 485, 501. In contrast, the amended version requires that the EPA grant the waiver “if the State determines that the State standards will be, in the aggregate, at least as protective of public health and welfare as applicable Federal standards.”232Clean Air Act Amendments of 1977, Pub. L. No. 95-95, § 207, 91 Stat. 685, 755 (emphasis added); see infra Section I.A. Congress intentionally granted California deference in creating motor vehicle standards in order to “afford California the broadest possible discretion in selecting the best means to protect the health of its citizens and the public welfare.”233H.R. Rep. No. 95-294, at 301–02 (1977); see MEMA I, 627 F. 2d at 1110–11. The amendment “confers broad discretion” on California to “weigh the degree of health hazards from various pollutants and the degree of emission reduction achievable for various pollutants with various emission control technologies and standards.”234H.R. Rep. No. 95-294, at 23 (1977). Congress made clear that the EPA should defer to California’s policy decisions, unless they are overwhelmingly arbitrary and capricious: the EPA Administrator “is not to overturn California’s judgment lightly. Nor is he to substitute his judgment for that of the State. There must be clear and compelling evidence that the State acted unreasonably in evaluating the relative risks of various pollutants . . . .”235Id. at 302. The EPA recognized in its 2013 waiver decision that Congress allowed it only limited review based on the section 209(b)(1) criteria to “ensure that the federal government did not second-guess state policy choices.”2362013 Waiver Grant, 78 Fed. Reg. at 2115. As the EPA affirmed, “Congress recognized that California could serve as a pioneer and a laboratory for the nation in setting new motor vehicle emission standards.”237Id. at 2113. Thus, as long as the regulations protect the health of California residents, the EPA should defer to California on the scope of those regulations.

ii.  Allowing California Broad Discretion to Regulate GHG Emissions Is Sufficiently Related to Addressing the Public Health Threats from Motor Vehicle Pollution in California

In Shelby County, the Voting Rights Act coverage formula factored in states’ voting discrimination history, which consisted of specific, unchangeable factors.238The coverage formula established that if the state used a law like a literacy or character test to keep people from registering to vote as of November 1, 1964, and less than 50% of the eligible voting population was registered to vote on November 1, 1964 or voted in the presidential election of November 1964, then the state was subject to preclearance. Voting Rights Act of 1965, Pub. L. 89-110, § 4(b), 79 Stat. 437. In contrast, Congress noted that California’s circumstances can change: if California no longer faces “compelling and extraordinary” conditions, it can no longer establish its own standards.239S. Rep. No. 90-403, at 33 (1967). This possibility creates a built-in mechanism to continually evaluate whether California needs its separate regulations240See Final Brief for Respondents, supra note 131, at 42. and whether the waiver provision is “justified by current needs.”241See Nw. Austin Mun. Util. Dist. No. 1 v. Holder, 557 U.S. 193, 203 (2009). Recognizing “the unique problems facing California as a result of its climate and topography,” Congress noted in 1967 that only California has demonstrated “compelling and extraordinary circumstances sufficiently different from the Nation as a whole to justify standards on automobile emissions which may, from time to time, need [to] be more stringent than national standards.”242H.R. Rep. No. 90-728, at 21–22 (1967); S. Rep. No. 90-403 at 33 (1967). The petitioners in Ohio v. EPA treat GHG emissions as if they are a separate and mutually exclusive concept from smog and criteria pollutants, claiming that because California has shifted from regulations to reduce local smog problems to regulations to reduce GHGs and address global climate change, the waiver provision no longer justifies California’s exemption.243See Brief for Petitioners, supra note 19, at 32 (“[T]here is no evidence California will suffer effects that are worse—in magnitude or in kind—than those experienced by the other forty-nine States.”). On the contrary, this Section argues that the effects of GHG emissions and smog pollution are interrelated and affect one another. Thus, addressing GHG emissions is directly related to Congress’s goal of addressing the public health threats from motor vehicle pollution in California.

Given the history of California’s early motor vehicle regulations and Congress’s interest in having California as a “laboratory for innovation” while not overburdening automobile manufacturers by forcing them to comply with multiple state standards, Congress intentionally struck a balance by authorizing just two standards: the national standard and the California standard.2442022 Waiver Reconsideration, 87 Fed. Reg. at 14360, 14377; H.R. Rep. No. 90-728, at 21 (1967); see S. Rep. No. 90-403 at 33–34 (1967). This compromise would allow California to continue to innovate and improve its air quality without creating a practical nightmare for automakers and interstate commerce.245Members of Congress favored states’ rights but were also concerned that having 50 different sets of requirements related to emissions controls would “unduly burden interstate commerce.” H.R. Rep. No. 95-294, at 309 (1977). Congress deliberately exempted California from federal preemption of motor vehicle regulations because of its “pioneering role in regulating automobile-related emissions, which pre-dated the Federal effort.”246Id. at 301. Because California had already adopted a robust air quality program and established its own motor vehicle emission standards prior to the passage of the federal Clean Air Act, it had expertise in emissions regulations that other states did not have.247See Ann E. Carlson, Federalism, Preemption, and Greenhouse Gas Emissions, 37 U.C. Davis L. Rev. 281, 314 (2003) (“The prospect of fifty separate standards for automobiles is untenable. But California has unique air pollution problems and an economy large enough to support separate standards.”); id. at 311 (noting that California “is probably unique in the country in the amount of expertise and sophistication it has developed in the regulation of auto emissions”).

California’s large automobile market and economy continue to justify its disparate treatment. At the time Congress passed the Clean Air Act waiver, it recognized the “presence and growth of California’s vehicle population, whose emissions were thought to be responsible for ninety percent of the air pollution in certain parts of California.”2482013 Waiver Grant, 78 Fed. Reg. at 2126. Congress noted the large effect of vehicles on local air pollution: “Motor vehicles are responsible for about 90 percent of the smog in the Los Angeles County, some 56 percent in the San Francisco Bay area, and about 50 percent in San Diego.”249H.R. Rep. No. 90-728, at 97 (1967). Congress also noted that because of its large size, California has “an economy large enough to support separate standards.”250Carlson, supra note 247, at 314. Thus, California’s market was large enough that automobile companies could still make a sizable profit while producing cars to meet California’s more stringent environmental requirements.251“The auto industry has shown itself willing and able to make the modifications required for its lucrative California market.” H.R. Rep. No. 90-728, at 97 (1967). There were twice as many vehicles in California as in any other state, including New York.252113 Cong. Rec. H30942 (daily ed. Nov. 2, 1967) (statement of Rep. Chet Holifield, California). Today, California continues to be the largest automobile market in the United States; if the state were a country, it would be the tenth largest auto market in the world.253Based on new passenger car/light vehicle registrations. Felix Richter, California Is Among the World’s Largest Car Markets, Statista (Sept. 24, 2020), https://www.statista.com/chart/23023/top-10-markets-for-new-passenger-car-registrations [https://perma.cc/6886-64FN]. California makes up 11% of U.S. new light-duty vehicle sales, and combined with the states that have already adopted its LEV rules, makes up 40.1% of U.S. new light-duty vehicle sales.254Cal. Air Res. Bd., supra note 8. Forty-three percent of ZEVs sold in the U.S. are sold in California.255California ZEV Sales Near 18% of All New Car Sales in 2022, Off. Cal. Governor Gavin Newsom (Oct. 19, 2022), https://www.gov.ca.gov/2022/10/19/california-zev-sales-near-18-of-all-new-car-sales-in-2022 [https://perma.cc/XM2W-6F3U].

California’s unique topography and climate conditions have also contributed to the air pollution problems exacerbated by climate change. The legislative history indicates that Congress granted California an exemption to regulate motor vehicle emissions primarily because California was facing unique, severe air pollution problems across the state, particularly in the Los Angeles area.256See H.R. Rep. No. 90-728, at 50 (1967) (recognizing the “critical concern of California for air pollution control, which is prompted especially by the acute susceptibility of the Los Angeles basin to concentrations of smog”). California’s air pollution problem was among “the most pervasive and acute in the Nation” at the time.257H.R. Rep. No. 95-294, at 301 (1977); see 113 Cong. Rec. H30943 (daily ed. Nov. 2, 1967) (statement of Rep. Tunney, California: “We are facing a serious and spreading smog problem, primarily caused by motor vehicle emissions.”). Geographical and climatic factors were consistently cited as “compelling and extraordinary” factors during the House debate, including the “unique problems facing California as the result of numerous thermal inversions that occur within that State because of its geography and prevailing winds pattern.”258113 Cong. Rec. H30948 (daily ed. Nov. 2, 1967) (statement of Rep. Harley Staggers, Chairman, House Interstate and Foreign Commerce Committee); see also id. at H30955 (statement of Rep. Roybal, California, referring to “atmospheric inversion”); id. at H30975 (statement of Rep. John Moss, California, referring to California’s “unique” meteorological problems). Rep. Holifield noted that California has a unique problem due to an atmospheric inversion which “the peculiar topography of the metropolitan area of Los Angeles County” has caused to some extent by keeping smog in the area and surrounding counties.259Id. at H30942 (statement of Rep. Chet Holifield, California). Even though members of Congress recognized that air pollution also affects other states in concerning ways,260William Macomber, Jr., Assistant Secretary for Congressional Relations, noted that air pollution has become an increasingly pressing problem in most metropolitan areas, including New York City, Detroit, Pittsburgh, Chicago, Baltimore, and Washington D.C. H.R. Rep. No. 90-728, at 50 (1967). they agreed that California’s distinct conditions and topography continue to contribute to the unique effects of pollution in the state, creating a critical need for air pollution control.261See S. Rep. No. 90-403, at 33 (1967) (“California’s unique problems and pioneering efforts justified a waiver . . . in the 15 years that auto emission standards have been debated and discussed, only the State of California has demonstrated compelling and extraordinary circumstances sufficiently different from the Nation as a whole . . . .”). As CARB established, California’s ozone levels will be exacerbated by higher temperatures from global warming, and “there is general consensus that temperature increases from climate change will exacerbate the historic climate, topography, and population factors conducive to smog formation in California, which were the driving forces behind Congress’s inclusion of the waiver provision.”2622022 Waiver Reconsideration, 87 Fed. Reg. at 14364 n.297.

Most significantly, climate change has only exacerbated the air pollution and smog problems that initially compelled California’s motor vehicle regulations and the Clean Air Act waiver. Automobiles emit both GHGs and smog-forming emissions including nitrogen oxide, carbon monoxide, and particulate matter.263Greenhouse Gas Versus Smog Forming Emissions, EPA, https://19january2017snapshot.epa.gov/greenvehicles/greenhouse-gas-versus-smog-forming-emissions_.html [https://perma.cc/ULA6-84AC]. The 2021 report of the United Nations’ Intergovernmental Panel on Climate Change (“IPCC”) reflects the latest scientific consensus that climate change is both a local and global problem.264Summary for Policymakers, in Climate Change 2021: The Physical Science Basis. Contribution of Working Group I to the Sixth Assessment Report of the Intergovernmental Panel on Climate Change 25 (Valérie Masson-Delmotte et al. eds., 2021) [hereinafter IPCC 2021 Report] (“Cities intensify human-induced warming locally, and further urbanization together with more frequent hot extremes will increase the severity of heatwaves.”). The report establishes a connection between climate change and intensifying weather extremes including heat waves and droughts.265Id. at 8. Additionally, GHGs contribute to respiratory disease from smog and air pollution.266Christina Nunez, Carbon Dioxide Levels are at a Record High. Here’s What You Need to Know, National Geographic (May 13, 2019), https://www.nationalgeographic.com/environment/article/greenhouse-gases [https://perma.cc/T2EQ-QABH]. GHG emissions lead to hotter global temperatures,267IPCC 2021 Report, supra note 264, at 5. which is expected to enhance the formation of ground-level ozone (a main component of smog).268John H. Tibbetts, Air Quality and Climate Change: A Delicate Balance, 123 Env’t Health Persps. A148, A149 (2015); Junfeng (Jim) Zhang, Yongjie Wei & Zhangfu Fang, Ozone Pollution: A Major Health Hazard Worldwide, 10 Frontiers Immunology 1, 2–3 (2019); Criteria Pollutants, N.H. Dep’t Env’t Servs., https://www.des.nh.gov/air/state-implementation-plans/criteria-pollutants [https://perma.cc/F8HD-GUFC] (noting ozone is a key ingredient in smog). Exposure to ozone can cause respiratory problems269Tibbetts, supra note 268, at A151. and aggravate lung diseases including asthma, particularly within more vulnerable groups.270Greenhouse Gas Versus Smog Forming Emissions, EPA, supra note 263; Health Effects of Ozone Pollution, EPA, https://www.epa.gov/ground-level-ozone-pollution/health-effects-ozone-pollution [https://perma.cc/LVH2-6KX8]; see also Ozone Effects, Cal. Air Res. Bd. (Nov. 3, 2016), https://ww2.arb.ca.gov/resources/fact-sheets/ozone-effects [https://perma.cc/P7TL-JJ4V]; Ozone and Your Health, Ctrs. for Disease Control & Prevention (Feb. 16, 2023), https://www.cdc.gov/air/ozone.html [https://perma.cc/YEB4-Z7XM]. Thus, GHGs can worsen exposure to ground-level ozone and smog, which is associated with increased mortality from respiratory and cardiovascular diseases.271Zhang et al., supra note 268, at 5. As a result, it has been well established that GHGs and smog are interrelated and affect air quality separately and together.272See 2022 Waiver Reconsideration, supra note 86, at 14363 (“[A]ir pollution problems, including local or regional air pollution problems, do not occur in isolation.”); see also Final Brief for Respondents, supra note 131, at 89–90.

Contrary to what the petitioners claim, climate change continues to uniquely affect California as an “acute California problem.”273See Final Brief for Respondents, supra note 131, at 52. While GHG emissions from California cars can “become one part of the global pool of GHG emissions,”2742008 Waiver Denial, 73 Fed. Reg. at 12160. this global pool eventually affects local conditions. The EPA recognized CARB’s strong evidence that California is “particularly impacted by climate change, including increasing risks from record-setting fires, heat waves, storm surges, sea-level rise, water supply shortages and extreme heat,” and that “GHG emissions contribute to local air pollution.”2752022 Waiver Reconsideration, 87 Fed. Reg. at 14363, 14365. Climate change impacts ozone exacerbation and wildfires, which affect local air quality.276Id. at 14334 n.10. California continues to have a serious smog problem, exacerbated by climate change.277California & the Waiver: The Facts, Cal. Air Res. Bd. (Sept. 17, 2019), https://ww2.arb.ca.gov/resources/fact-sheets/california-waiver-facts [https://perma.cc/N9DL-6B2P]. Seven of the ten cities with the worst air pollution nationwide are in California.278Id.; see Most Polluted Cities, Am. Lung Ass’n, https://www.lung.org/research/sota/city-rankings/most-polluted-cities [https://perma.cc/Z535-6KNT]. Ten million Californians in the San Joaquin Valley and Los Angeles air basins currently live under “severe non-attainment” conditions for ozone, where people suffer unusually high rates of asthma and cardiopulmonary disease.279Cal. Air Res. Bd. supra note 277. Climate change has increased the number of hot days that can result in smog events and exacerbate wildfires.280Id. Thus, smog exacerbates climate change, which in turn exacerbates smog, and GHGs—which lead to climate change—continue to pose a direct and local threat.281Cause and Effects of Climate Change, U.N., https://www.un.org/en/climatechange/science/causes-effects-climate-change [https://perma.cc/6G32-UAYX] (“As greenhouse gas emissions blanket the Earth, they trap the sun’s heat. This leads to global warming and climate change.”). As the 2022 EPA decision concluded, the 2019 EPA decision to withdraw the 2013 EPA waiver grant failed to properly consider “the nature and magnitude of California’s serious air quality problems, including the interrelationship between criteria and GHG pollution.”2822022 Waiver Reconsideration, 87 Fed. Reg. at 14334. The EPA noted that the 2019 record contained evidence that GHG emissions can lead to locally elevated carbon dioxide concentrations with local impacts such as ocean acidification, in addition to the longer-term global impacts from global emissions.283Id. at 14366. Thus, just like smog, climate change poses serious threats to the public health and safety of residents in California. As a result, ZEV regulations are crucial in protecting the public health and safety of Californians.

Even adopting the 2019 EPA’s narrow “local nexus” test, which required that the California waiver only applies to measures that address conditions “extraordinary” with respect to California, or those with a specific connection to local features and emissions peculiar to California,2842019 Waiver Withdrawal, 84 Fed. Reg. at 51347. California’s ZEV standard meets this test in directly addressing local air pollutant conditions by reducing criteria pollutant emissions. California’s 2020 Executive Order and resulting ACC II regulations made clear that California intended to regulate both GHG emissions and smog pollutants. The 2020 Executive Order states that zero emissions technologies “reduce both greenhouse gas emissions and toxic air pollutants,”285Cal. Exec. Order No. N-79-20 (Sept. 23, 2020), https://www.gov.ca.gov/wp-content/uploads/2020/09/9.23.20-EO-N-79-20-Climate.pdf [https://perma.cc/F4SE-B5AB]. and the ACC II regulations require new vehicles to “produce zero exhaust emissions of any criteria pollutant (or precursor pollutant) or greenhouse gas . . . .”286Cal. Code Regs. tit. 13, § 1962.4. California’s more stringent standards will thus continue to achieve critical reductions in conventional criteria pollution and help the state address public health problems caused by smog and soot.287See 2022 Waiver Reconsideration, 87 Fed. Reg. at 14353 (“CARB’s motor vehicle emission standards operate in tandem and are designed to reduce both criteria and GHG pollution and the ways in which GHG pollution exacerbates California’s serious air quality problems, including the heat exacerbation of ozone . . . .”); id. at 14364 (“CARB had demonstrated the need for GHG standards to address criteria pollutant concentrations in California.”). Congress has not provided any indication that California cannot take measures to reduce criteria pollutants and GHGs. Transportation is the largest source of air pollution in the state, responsible for nearly 40% of GHG emissions, 80% of nitrogen oxide pollution, and 90% of diesel particulate matter pollution.288Transforming Transportation, Cal. Energy Comm’n, supra note 2; Current California GHG Emission Inventory Data, Cal. Air Res. Bd., supra note 2. The EPA concluded that GHG measures are relevant to addressing local criteria pollutant issues2892009 Waiver Grant, supra note 68, at 32763 (“[A]lthough the factors that cause ozone are primarily local in nature and [] ozone is a local or regional air pollution problem, the impacts of global climate change can nevertheless exacerbate this local air pollution problem . . . California has made a case that its greenhouse gas standards are linked to amelioration of California’s smog problems. Reducing ozone levels in California cities and agricultural areas is expected to become harder with advancing climate change . . . ‘California’s high ozone levels—clearly a condition Congress considered—will be exacerbated by higher temperatures from global warming.’ ”); id. at 32750 (“CARB also found that its greenhouse gas standards will increase the health and welfare benefits from its broader motor vehicle emissions program by directly reducing upstream emissions of criteria pollutants from decreased fuel consumption.”). and that regulations to reduce GHGs often simultaneously address smog-forming pollutants like nitrogen oxide.2902022 Waiver Reconsideration, 87 Fed. Reg. at 14364 (citing Heavy-Duty Tractor-Trailer Greenhouse Gas Regulations); Notice of Decision, 79 Fed. Reg. 46256, 46261 (Aug. 7, 2014) (projecting that GHG standards will reduce nitrogen oxide emissions by one to three tons per day through 2020). The legislative history provides no basis for the claim that California cannot mitigate climate change threats or address environmental problems within their boundaries as soon as the problems extend beyond them.291See Final Brief for Respondents, supra note 131, at 52. In fact, Congress expressed an interest in allowing California to “continue its already excellent program” and continue to be the testing area of motor vehicle standards, which is expected to benefit its people and the nation by strengthening federal standards.292S. Rep. No. 90-403, at 33 (1967). The Senate report reflected opposition to displacing California’s right to set more stringent standards, as justified by California’s “unique problems and pioneering efforts.”293Id. Members of Congress concurred with the principle that California’s advances in air pollution regulation should not be nullified and that the state’s progress should not be impeded. Congressman John Dingell stated: “To penalize California for being ahead of the rest of the country in combating the menace of air pollution is totally incomprehensible.”294113 Cong. Rec. at H30946 (daily ed. Nov. 2, 1967) (remarks of Congressman John Dingell). The Ninth Circuit has also stated that California should be “encouraged to continue and to expand its efforts . . . to lower carbon emissions.”295Rocky Mountain Farmers Union v. Corey, 730 F.3d 1070, 1107 (9th Cir. 2013). Thus, Congress’s reasons for granting California a waiver continue to be compelling and extraordinary, and California’s current needs continue to remain relevant as ever in justifying the Clean Air Act waiver provision.

Congress did not justify the Clean Air Act waiver provision based on whether pollution problems were of a more local or global nature, but rather on the unique effects of smog in the Los Angeles area.296See H.R. Rep. No. 90-728, at 50 (1967) (recognizing the “critical concern of California for air pollution control, which is prompted especially by the acute susceptibility of the Los Angeles basin to concentrations of smog”). This emphasis suggests that Congress intended to give California the flexibility to adopt motor vehicle standards that the state determines are needed to address air pollution in the state, regardless of whether those problems might also be global in nature.297See 2022 Waiver Reconsideration, 87 Fed. Reg. at 14363 (“EPA sees no reason to distinguish between ‘local or regional’ air pollutants versus other pollutants that may be more globally mixed. Rather, it is appropriate to acknowledge that all pollutants and their effects may play a role in creating air pollution problems in California and that EPA should provide deference to California in its comprehensive policy choices for addressing them.”). Thus, California’s problems are serious enough and its efforts are such a model for the nation that a waiver provision is necessary in order for California to adequately protect public health. More recently, Congress’s clarification in the 2022 Inflation Reduction Act that GHGs are pollutants regulated under the Clean Air Act suggests that Congress intends the Clean Air Act to include GHGs.298Inflation Reduction Act of 2022, Pub. L. No. 117-169, 136 Stat. 1818. This further strengthens the argument that California is acting within the scope of the Clean Air Act in regulating GHGs through its innovative motor vehicle program.

CONCLUSION

The equal sovereignty argument is a new attempt to invalidate the Clean Air Act waiver provision and California’s ability to regulate motor vehicle emissions. As of this Note, no court has specifically addressed the constitutionality of the Clean Air Act under the equal sovereignty principle, and the decision is pending for Ohio v. EPA, which is expected to address this constitutional question.

This Note concludes that the equal sovereignty principle does not apply to the Clean Air Act, but even if it were to apply, it does not invalidate section 209(b)(1). Distinguishing from the outcome in Shelby County, the Clean Air Act waiver provision remains constitutional because granting California an exemption is “sufficiently related to the problem that it targets.” First, the Clean Air Act targets the broader problem of public health from automobile emissions. Second, allowing California to implement more stringent motor vehicle regulations will directly help address this problem. Congress had strong justifications for granting California an exemption which continue to remain compelling and relevant today. California’s history with air pollution control, its large economy, and its characteristic geographic and climate conditions put the state in a unique position to influence the automobile market and address GHG emissions. California faces new and increasingly formidable threats from climate change, which have exacerbated the existing problems that initially compelled California’s motor vehicle regulations. Allowing California broad discretion to regulate GHG emissions is directly related to Congress’s goal of addressing the public health threats from motor vehicle pollution in California because the effects of GHGs and smog are directly related and affect one another. Even as California’s motor vehicle regulations have shifted from reducing local smog by regulating criteria pollutants to reducing GHG emissions by eliminating gasoline-powered cars, California’s current needs continue to justify its differential treatment—maintaining, and perhaps even strengthening, section 209(b)(1)’s relevance in the twenty-first century.

The court’s decision on whether section 209(b)(1) of the Clean Air Act remains constitutionally valid will determine the extent to which California can continue to realize the localized benefits of the Clean Air Act while helping accelerate the nation’s transition towards a clean energy economy. It will also have implications for California’s ability to continue to regulate GHG emissions as a leader in addressing the most pressing environmental issues of the day.

Is the court going to handcuff California’s ability to protect the health and safety of its residents in the name of equal sovereignty? That was not the intention of Congress when it discussed equal sovereignty concerns pertaining to the Clean Air Act waiver. On the contrary, Congress debated whether other states should also be able to enact more stringent standards than the federal government, which would be the more reasonable remedy if the Clean Air Act waiver provision were deemed unconstitutional per equal sovereignty, as the petitioners demand.

To strengthen the ability of motor vehicle regulations to withstand future court challenges, California could emphasize criteria pollutants in its regulations. Since criteria pollutants have been more directly linked to local air pollution issues and Congress originally implemented the waiver provision in response to regional smog problems, this change could make it more difficult to challenge a regulation on the basis of it only regulating climate change. It will likely be simpler to show that the disparate treatment of California is sufficiently related to the problem that the Clean Air Act targets if legislators explicitly provide how they expect the regulations to affect local air quality as well as the local co-benefits of implementing them. For example, replacing internal combustion passenger vehicles with EVs will reduce not only GHG emissions, but also criteria pollutants including nitrogen oxides that are emitted.

California’s motor vehicle standards alone may not reverse or solve climate change, but the EPA has a duty to take steps to slow or reduce it.299States need not “resolve massive problems in one fell regulatory swoop.” Massachusetts v. EPA, 549 U.S. 497, 524 (2007). Allowing California to continue to promulgate innovative, forward-looking motor vehicle standards is crucial to its ability to lead the country as a “laboratory of innovation,” as Congress intended, and address the urgent environment and public health consequences of motor vehicle pollution.

97 S. Cal. L. Rev. 165

Download

* Senior Editor, Southern California Law Review, Volume 97; J.D. Candidate 2024, University of Southern California Gould School of Law; B.A. Economics 2019, Wellesley College. A special thank you to Professor Robin Craig for her thoughtful guidance, my friends and family for their consistent support and encouragement, and the Southern California Law Review editors for their thorough feedback.

Judging Firearms Evidence

Firearms violence results in hundreds of thousands of criminal investigations each year. To try to identify a culprit, firearms examiners seek to link fired shell casings or bullets from crime scene evidence to a particular firearm. The underlying assumption is that firearms impart unique marks on bullets and cartridge cases, and that trained examiners can identify these marks to determine which were fired by the same gun. For over a hundred years, firearms examiners have testified that they can conclusively identify the source of a bullet or cartridge case. In recent years, however, research scientists have called into question the validity and reliability of such testimony. Judges largely did not view such testimony with increased skepticism after the Supreme Court set out standards for screening expert evidence in Daubert v. Merrell Dow Pharmaceuticals, Inc. Instead, the surge in judicial rulings came more than a decade later, particularly after reports by scientists shed light on limitations of the evidence.

In this Article, we detail over a century of case law and examine how judges have engaged with the changing practice and scientific understanding of firearms comparison evidence. We first describe how judges initially viewed firearms comparison evidence skeptically and thought jurors capable of making firearms comparisons themselves—without an expert. Next, judges embraced the testimony of experts who offered more specific and aggressive claims, and the work spread nationally. Finally, we explore the modern era of firearms case law and research. Judges increasingly express skepticism and adopt a range of approaches to limit in-court testimony by firearms examiners.

In December 2023, Rule 702 of the Federal Rules of Evidence was amended, for the first time in over twenty years, specifically due to the Rules Committee’s concern with the quality of federal rulings regarding forensic evidence, as well as the failure to engage with the ways that forensic experts express conclusions in court. There is perhaps no area in which judges, especially federal judges, have been more active than in the area of firearms evidence. Thus, the judging of firearms evidence has central significance for the direction that scientific evidence gatekeeping may take under the revised Rule 702 in federal, and then state courts. We conclude by examining lessons regarding the gradual judicial shift toward a more scientific approach. The more-than-a-century-long arc of judicial review of firearms evidence in the United States suggests that, over time, scientific research can displace tradition and precedent to improve the quality of justice.

INTRODUCTION

On November 11, 2016, a police officer recovered a forty-caliber Smith & Wesson cartridge casing from the scene of a homicide in Washington D.C.1See United States v. Tibbs, No. 2016 CF1 19431, 2019 D.C. Super. LEXIS 9, at *8 (D.C. Super. Ct. Sept. 5, 2019). A police officer reported seeing a person discarding a Smith & Wesson semiautomatic pistol shortly after the homicide occurred.2Id. Police sent a recovered cartridge casing to the crime lab where an examiner identified it—conclusively—“as having been fired” by the pistol recovered from the defendant,3Id. at *8–9. charged with first-degree murder.4Id. at *8. As the case approached trial, the defense challenged the admissibility of this proffered expert testimony, arguing it should be excluded because it was not the “product of reliable principles and methods.”5Id. at *12. One of the authors served as an expert in the case. See id. at *9. In other words, the method lacked “scientific validity.” After hearing from several experts and reviewing published studies, Washington D.C. Superior Court Associate Judge Edelman found that there was insufficient evidence that firearms examiners can reliably make an identification.6Id. at *3 (“According to the government’s proffer, this analysis permitted the examiner to identify the recovered firearm as the source of the cartridge casing collected from the scene.”). The judge ruled an expert could—at most—opine that “the recovered firearm cannot be excluded as the source of the cartridge casing found on the scene of the alleged shooting.”7Id. at *77 (emphasis added); see also id. at *2. As we will describe, this is a powerful new limit on firearms evidence, a field in which experts have confidently concluded for decades that one and only one firearm—to the exclusion of all other firearms in the world—can produce the ammunition found at a given crime scene.8See Brandon L. Garrett, Nicholas Scurich & William E. Crozier, Mock Jurors’ Evaluation of Firearm Examiner Testimony, 44 Law & Hum. Behav. 412, 413 (2020) (studying jury evaluation of firearm expert testimony and finding “cannot exclude” language to influence verdicts); infra Part II.

While this case represented just one trial judge’s ruling, it not only forms a part of a sea change in judicial review of firearms evidence, but also the local repercussions point to more fundamental problems in our criminal system. Consider a later case before Judge Edelman, this one with charges brought against two men for two killings involving firearms evidence. Prosecutors were understandably concerned.9Jack Moore, DC Judge Orders Forensic Lab to Turn Over Some Documents Sought by Prosecutors, WTOP News (Nov. 10, 2020, 2:34 PM), https://wtop.com/dc/2020/11/dc-judge-orders-forensic-lab-to-turn-over-some-documents-sought-by-prosecutors [https://perma.cc/L5X2-JGJG]. In this case, D.C.’s Metropolitan Crime Lab had reported that the same weapon fired the cartridge casings found at each crime scene.10Id. Perhaps because they feared that the judge might view the evidence with renewed skepticism, the prosecutors took an unusual step: they asked independent examiners to take a look at the evidence.11Id.

The independent experts definitively concluded that two different firearms were involved—the opposite of what the D.C. crime lab examiners had concluded.12Jack Moore & Megan Cloherty, ‘You Can Trust This Laboratory’: DC Crime Lab Director Responds to Scrutiny of Firearms Unit, WTOP News (Dec. 2, 2020, 4:24 AM), https://wtop.com/dc/2020
/12/you-can-trust-this-laboratory-dc-crime-lab-director-responds-to-scrutiny-of-firearms-unit [https://
perma.cc/X93R-SUP7].
Internally, the lab examiners reexamined the evidence and agreed the cartridges came from different weapons. After meeting with lab managers, however, they instead reported an altered finding of “inconclusive,” meaning that no conclusion could be reached.13Prosecution’s Praecipe at 3, United States v. McLeod, No. 2017-CF-19869 (D.C. Super. Ct. Mar. 22, 2021). The management notified the ANSI National Accreditation Board (“ANAB”), which accredited the lab, that an internal review resulted in an “inconclusive” finding, but the audit that followed found that the lab managers had acted to conceal the errors in the case.14See id. at 2–3 (“DFS management not only failed to properly address the conflicting results reported to the DFS by the USAO, but also engaged in actions to alter the results reached by the examiners assigned to conduct a reexamination of the evidence.”). In April 2020, ANAB suspended the lab’s accreditation, and as a result, the lab was shut down.15Keith L. Alexander, National Forensics Board Suspends D.C. Crime Lab’s Accreditation, Halting Analysis of Evidence, City Says, Wash. Post, (Apr. 3, 2021, 7:43 PM), https://www.washington
post.com/local/public-safety/dc-lab-forensic-evidence-accreditation/2021/04/03/723c4832-94aa-11eb-a74e-1f4cf89fd948_story.html [https://perma.cc/2YS5-Y6QG].
Prosecutors then opened a new probe into its firearms unit, the lab director resigned,16Paul Wagner, D.C. Crime Lab Under Investigation After Allegations of Wrongdoing, NBC News (Apr. 8, 2021, 8:40 PM), https://www.nbcwashington.com/news/local/dc-crime-lab-under-investi
gation-after-allegations-of-wrongdoing/2634489 [https://perma.cc/4NJ5-GP4K].
the lab disbanded, and the firearms unit remains closed as of this writing.17Jack Moore, D.C. Abruptly Disbands Crime Lab’s Firearms Unit, WTOP News (Sept. 16, 2021, 4:00 PM), https://wtop.com/dc/2021/09/dc-abruptly-disbands-crime-labs-firearms-unit [https://
perma.cc/C3YN-LCYJ]. It appears that in December 2023, the D.C. crime lab regained partial accreditation.  As of this writing, however, the firearms unit has not regained accreditation, and it remains closed. Mark Segraves, DC Forensic Crime Labs Regain Accreditation After Nearly 3 Years, NBC Wash. (Dec. 27, 2023, 1:25 PM), https://www.nbcwashington.com/news/local/dc-forensic-crime-labs-regain-accreditation-after-nearly-3-years/3501258 [https://perma.cc/U342-NCE5]; Ivy Lyons, DC Crime Lab Appears to Regain Partial Accreditation After Losing Ability to Process Evidence in 2021, WTOP News (Dec. 26, 2023, 3:11 PM), https://wtop.com/dc/2023/12/dc-crime-lab-regains-some-accreditation-3-years-after-losing-ability-to-process-evidence [https://perma.cc/2TGY-USKX].

This rapidly unfolding crisis began with a spot-check in a single case prompted by a judge asking a fundamental question: How often do firearms examiners get it right versus wrong? For decades, few judges asked the question, but as we detail in this Article, judges have become increasingly engaged with the underlying science and have transformed a backwater area of forensic evidence into a subject of complex litigation. Indeed, in no other area have judges engaged in such a detailed manner with the limits of the testimony expressed by examiners—making firearms evidence the most prominent testing ground for the 2023 amendments to the Federal Rules of Evidence, designed to tighten judicial review of experts more generally, but with a focus on forensic evidence more specifically.18Advisory Comm. on Rules of Prac. and Proc., June 2022 Agenda Book 891–93 (2022) [hereinafter 2022 Comm. on Rules of Prac. and Proc.]; Fed. R. Evid. 702 (2023 amendment).

Firearms examination is in great demand, with more than a hundred thousand requests for a forensic firearm examination each year in the United States.19See Matthew R. Durose, Andrea M. Burch, Kelly Walsh & Emily Tiry, Bureau of Just. Stats., NCJ 250151, Publicly Funded Forensic Crime Laboratories: Resources and Services, 2014 3 (2016). Firearms violence is a major problem in the United States—more than ten thousand homicides and almost five hundred thousand other crimes, such as robberies and assaults, are committed using firearms.20See Gun Violence in America, Nat’l Inst. of Just. (Feb. 26, 2019), https://www.nij.gov/
topics/crime/gun-violence/pages/welcome.aspx [https://perma.cc/4TXL-K3NC]; 2018 January-June Preliminary Semiannual Uniform Crime Report: Crime in the United States, FBI (2018), https://ucr.fbi.
gov/crime-in-the-u.s/2018/preliminary-report [https://perma.cc/VMU8-ZYSG].
When conducting these comparisons, examiners seek to link crime scene evidence—such as spent cartridge casings or bullets—with a firearm. These examiners assume that the manufacturing processes used to cut, drill, and grind a gun leaves distinct and identifiable markings on the gun’s barrel, breech face, firing pin, and other components. When the firearm discharges, those components in turn contact the ammunition and leave marks on it. Experts have long assumed, as we will describe, that firearms leave distinct toolmarks on ammunition.21See infra Section I.A. They believe that they can definitively link spent ammunition to a particular firearm using these toolmarks.22See id. And for over a hundred years, examiners have offered criminal trial testimony relying on this assumption.23See infra Part I.

In recent years, the consequences of the uncritical judicial acceptance of firearms comparison testimony have come into sharper focus. Indeed, we now know that firearms evidence played a central role in numerous high-profile wrongful convictions. In the 2014 per curiam opinion in Hinton v. Alabama, for example, the U.S. Supreme Court reversed a conviction due to the defense lawyer’s inadequate performance in failing to develop firearms evidence at a capital murder trial.24Hinton v. Alabama, 571 U.S. 263, 264 (2014). The central evidence was a State Department of Forensic Sciences examiner’s conclusion that six bullets were fired from the same gun: “[T]he revolver found at Hinton’s house.”25Id. at 265. The defense did not hire a competent and qualified expert, and the Court emphasized that “the only reasonable and available defense strategy require[d] consultation with experts or introduction of expert evidence.”26Id. at 273 (quoting Harrington v. Richter, 562 U.S. 86, 106 (2011)). Hinton was subsequently exonerated, and he commented: “I shouldn’t have [sat] on death row for thirty years . . . . All they had to do was to test the gun.”27Abby Phillip, Alabama Inmate Free After Three Decades on Death Row: How the Case Against Him Unraveled, Wash. Post (Apr. 3, 2015, 10:28 PM), https://www.washingtonpost.com/
news/morning-mix/wp/2015/04/03/how-the-case-against-anthony-hinton-on-death-row-for-30-years-unraveled [https://perma.cc/5QPA-4M83].

This Article presents the results of a comprehensive review of all judicial rulings in the United States concerning firearms comparison evidence. Our database of more than 300 judicial rulings is available as a resource online.28See Firearms Expert Evidence Database, Ctr. for Stats. and Applications in Forensic Evidence (2022), https://forensicstats.org/firearms-expert-evidence-database [https://perma.cc/LR4J-RLU4]. The database “ha[s] assembled reported decisions, chiefly by appellate courts, that discuss the admissibility of expert testimony regarding firearms comparison evidence.” Id. The database consists of written, published decisions (largely appellate opinions but also some trial rulings).29The cases that are included in this database were:

[G]athered using searches of the Westlaw legal database, across all fifty states and the federal government, with rulings dating back over one hundred years. Where possible, trial rulings were obtained, but generally these cases reflect reported, written decisions containing the keywords used, and therefore largely reflect appellate rulings. The cases are searchable across a range of characteristics, including basic information concerning the state, year, type of court, and parties, but also details concerning the basis of the rulings and the factors relied upon by each court. The database describes whether the ruling employed a Daubert or Frye standard, or a ruling regarding local rules of evidence, and what the result of that ruling was.

Id.
We describe the three-part story of the path of firearms evidence: (1) initial skepticism of a novel set of methods, then moving to; (2) national acceptance of increasingly powerfully stated conclusions regarding firearms; and finally (3) a surge in judicial opinions and skepticism of firearms comparison evidence that followed, not Daubert and the new reliability-focused standards for judicial review of scientific evidence, but rather a series of scathing reports by the scientific community calling into question the reliability of firearms evidence.

First, we describe how in the earliest cases, judges were actually quite skeptical of firearms comparison evidence, particularly when presented by self-styled experts, and often concluded that jurors were capable of making the comparisons themselves, without a need for expert testimony.30See infra Part I. However, particularly due to the influence of the flamboyant Major Calvin Goddard and his disciples, courts gradually embraced the firearms comparison evidence as the subject of expert testimony.31See infra Part I.

Second, we document how the claims made by experts became more specific and aggressive as the work spread nationally.32See infra Part I. Rather than simply describing a comparison between two sets of objects, firearms experts testified by making “uniqueness” claims: the theory that “no two firearms should produce the same microscopic features on bullets and cartridge cases such that they could be falsely identified as having been fired from the same firearm.”33Erich D. Smith, Cartridge Case and Bullet Comparison Validation Study with Firearms Submitted in Casework, 36 AFTE J. 130, 130 (2004) (quoted in United States v. Monteiro, 407 F. Supp. 2d 351, 361 (D. Mass. 2006)). By the 1960s, this expert testimony was offered and accepted across the country. Professional groups emerged and set standards for the field, which courts took note of. Written judicial opinions became quite uncommon, and any judicial skepticism was largely limited to more unusual applications of the methods rather than the underlying methodology itself.34See infra Part I.

Third, we explore the modern era of firearms case law and research, with increasingly intense judicial interest and written opinions on the topic in the last two decades.35See infra Part II. In 1993, the Supreme Court decided Daubert v. Merrell Dow Pharmaceuticals36Daubert v. Merrell Dow Pharm., Inc., 509 U.S. 579 (1993). and, along with its progeny and the revision to Federal Rule of Evidence 702 (“Rule 702”) and state-law analogues, judges now bear clearer and more rigorous gatekeeping responsibilities to assess the reliability of scientific evidence.37See generally, e.g., David L. Faigman, The Daubert Revolution and the Birth of Modernity: Managing Scientific Evidence in the Age of Science, 46 U.C. Davis L. Rev. 893 (2013). Accompanying this shift in the courts, by the late 1990s, experts premised testimony on a “theory of identification” set out by a professional association, the Association of Firearms and Tool Mark Examiners (“AFTE”).38See infra Part II. The AFTE instructs practitioners to use the phrase “source identification” to explain what they mean when they identify “sufficient agreement” of markings when examining bullets or cartridge cases.39What Is Firearm and Toolmark Identification?, The Ass’n of Firearm and Toolmark Examiners, https://afte.org/about-us/what-is-afte/what-is-firearm-and-tool-mark-identification [https://
perma.cc/XAU7-5Y4M].

In recent years, scientists have called into question the validity and reliability of this testimony—contributing to an explosion of judicial rulings. In a 2008 report, the National Academy of Sciences (“NAS”) found that “[t]he validity of the fundamental assumptions of uniqueness and reproducibility of firearms-related toolmarks has not yet been fully demonstrated.”40Nat’l Rsch. Council of the Nat’l Acads., Ballistic Imaging 81 (Daniel L. Cork et al. eds., 2008) [hereinafter 2008 NAS Report]. In its 2009 report, the NAS concluded “[s]ufficient studies have not been done to understand the reliability and repeatability of the methods.”41Nat’l Rsch. Council of the Nat’l Acads., Strengthening Forensic Science in the United States: A Path Forward 154 (2009) [hereinafter 2009 NAS Report]. The report also noted that “the lack of a precisely defined process . . . [that] does not even consider, let alone address, questions regarding variability, reliability, repeatability, or the number of correlations needed to achieve a given degree of confidence.”42Id. at 155. Judges have also raised concerns about the lack of specificity in the examination process. See, e.g., United States v. Green, 405 F. Supp. 2d 104, 114 (D. Mass. 2005) (stating the method is “either tautological or wholly subjective”); United States v. Shipp, 422 F. Supp. 3d 762, 779 (E.D.N.Y. 2019) (“[T]he sufficient agreement standard is circular and subjective.”). Over half of the judicial rulings that we identified have occurred since 2009, the year that the NAS issued its pathbreaking report. We detail dozens of opinions that have limited testimony of firearms experts in increasingly stringent ways.

Solidifying this trend, in 2016, the President’s Council of Advisors on Science and Technology (“PCAST”) reviewed in detail all of the firearm examiner studies that had been conducted to date,43President’s Council of Advisors on Sci. and Tech., Forensic Science in the Criminal Courts: Ensuring Scientific Validity of Feature-Comparison Methods X (Sept. 2016) [hereinafter PCAST Report]. finding, with only one deemed appropriately designed, that “the current evidence falls short of the scientific criteria for foundational validity.”44Id. at 111. Most recently—beginning in the aforementioned 2019 case before Judge Edelman—scientists have testified about the research base of firearm examination.45David L. Faigman, Nicholas Scurich & Thomas D. Albright, The Field of Firearms Forensics is Flawed, Sci. Am. (May 25, 2022), https://www.scientificamerican.com/article/the-field-of-firearms-forensics-is-flawed [https://perma.cc/ZM4A-TLMQ]. These experts include psychologists, statisticians, and other academics with training in conducting science, rather than applying a forensic technique. As one judge put it, “[R]arely do the experts fall into such cognizable camps, forensic practitioners on one side and academic researchers on the other.”46People v. Ross, 129 N.Y.S.3d 629, 639 (N.Y. Sup. Ct. 2020).

The impact of these modern critiques on the admissibility of firearm examination has borne concrete results, but gradually. Comforted by more than a century of long-standing precedent, judges were slow to react to scientific concerns raised regarding firearms comparison evidence, even after the Daubert ruling. Yet in more recent years, as lawyers have increasingly litigated the findings of scientific reports and error rate studies, we have seen a dramatic rise in a judge’s willingness to engage with scientific limitations of the methods.47See infra Section II.D. That said, most judges have responded by imposing limits on how experts phrase conclusions in testimony, but we note there are reasons to doubt that this compromise solution will sufficiently inform lay jurors of the limits of the method.48Regarding effectiveness of such measures, see Garrett et al., supra note 8, at 421–22. For further discussion, see infra Part III.

For the first time since 2000, Federal Rule of Evidence 702 was amended, as of December 1, 2023.492022 Comm. on Rules of Prac. and Proc., supra note 18, at 891–93. The Advisory Committee notes emphasize that these revisions are “especially pertinent” to forensic evidence.50Memorandum from the chair of the Committee on Rules of Practice and Procedure to the clerk of the Supreme Court 227 (Oct. 19, 2022), https://www.uscourts.gov/sites/default/files/2022_scotus_
package_0.pdf [https://perma.cc/QS33-9DTQ].
Further, for forensic pattern-comparison methods like firearms evidence, the committee noted that opinions “must be limited to those inferences that can reasonably be drawn from a reliable application of the principles and methods.”51Id. at 230. The amended Rule 702 specifically directs judges to (1) more carefully consider that the proponent of an expert bears the burden to show that the various reliability requirements are met and (2) underscore that the opinions that the expert formed are reliably supported by the application of the methods to the data.522022 Comm. on Rules of Prac. and Proc., supra note 18, at 891–93. The rule changes squarely address the issues that judges have grappled with in the area of firearms evidence, perhaps more prominently than in any other area of scientific evidence. The rule changes target the two main concerns that judges have raised: the reliability of the methods and the overstatement of conclusions.

Thus, the body of case law regarding firearms evidence may only grow, and it may be a harbinger for how judges will engage with scientific evidence more broadly after the rule change. In a 2023 ruling, the Supreme Court of Maryland ruled that an expert can only opine on whether spent bullets or cartridges are “consistent or inconsistent” with those known to have been fired by a particular weapon.53Abruquah v. State, 483 Md. 637, 648 (2023). In perhaps a sign of things to come, a trial judge in Cook County, Illinois recently excluded firearms expert testimony entirely, based on scientific concerns with reliability, after conducting an extensive evidentiary hearing. There, the judge concluded that the probative value of the evidence was a “big zero” and raised the concern of “yet another wrongful conviction” based on such evidence if the jurors viewed “[t]he combination of scary weapons, spent bullets, and death pictures without even a minimal connection” to expertise that is repeatable and reproducible.54See People v. Winfield, No. 15-CR-1406601, at 32–34 (Cir. Ct. Cook Cnty. Ill. Feb. 8, 2023).

These developments more fundamentally suggest that for judges and lawyers to carefully engage with the reliability rules set out in Daubert and in Rule 702, it takes engagement by the scientific community. Prominent scientific reports and studies have helped judges and lawyers apply scientific criteria to firearms examinations. The result has limited unsupported use of these firearms comparisons and may promote better methods in the future that can prevent errors and wrongful convictions.55See infra Section II.E. The changes to Rule 702 can cement these developments and ensure more careful review of scientific expert evidence more broadly. We conclude by examining the lessons to be learned from this more-than-a-century-long arc of judicial review of firearms evidence in the United States for future judicial engagement with science.

I.  FIREARMS METHODS AND THE FIRST HALF-CENTURY OF JUDICIAL RULINGS

In this Part, we begin by describing the basic approach used by firearms and toolmark examiners. The approach has been in use for over a hundred years, and its origins trace to a single pioneering examiner, Major Calvin H. Goddard, who powerfully transformed courts’ early skepticism toward firearms comparison evidence to near-universal acceptance.56Calvin Hooker Goddard—Father of Forensic Ballistics, Forensic’s Blog, https://forensicfield.blog/calvin-hooker-goddard-father-of-forensic-ballistics [https://perma.cc/69BV-KYQE] (last visited Sept. 22, 2023). Considered the “father” of modern forensic firearms examination, Goddard assembled databases of information from gun makers and pioneered a “comparison microscope,” a device with side-by-side eyepieces, to make comparing firearms evidence more convenient.57Id. While quite primitive compared with modern technology, Goddard introduced the use of the microscope in firearms comparison, which was seen as permitting a level of sophisticated visual analysis that a layperson lacked access to. We describe how, in the 1930s, Goddard often testified in trials about the comparison microscope, further cementing the method’s legitimacy to courts. Over time, other practitioners and crime laboratories adopted similar methods and began to testify as experts. We describe in this Part what reasoning courts used through the 1930s as they moved from early skepticism to acceptance of this expert testimony.

A.  A Primer on Firearm and Toolmark Identification

Toolmark identification is the practice of human observers opining on whether toolmarks were produced by a particular tool.58Id. A tool is considered any device that serves a mechanical purpose (for example, screwdrivers, pliers, knives, pipe wrenches). As the tool contacts softer material, it sometimes leaves marks on the softer object’s surface. The resulting marks are called “toolmarks.”59One text gives the following example: “For example, when a butter knife is dragged along the surface of butter, one may observe a series of lines across the top of the butter. In this case, the mark in the butter is a toolmark and the knife is the tool that made the mark.” Ronald Nichols, Firearm and Toolmark Identification: The Scientific Reliability of the Forensic Science Discipline 1 (2018). A firearm consists of many tools that perform mechanical functions to fire a bullet. Therefore, firearm identification is considered a subspecialty of toolmark identification.60United States v. McCluskey, No. 10-2734, 2013 U.S. Dist. LEXIS 203723, at *7 (D.N.M. Feb. 7, 2013) (“Firearm identification is a specialized area of toolmark identification dealing with firearms, which involve a specific category of tools.”). The goal of firearm identification is to determine whether two bullets or cartridge cases were fired by the same firearm.

Firearm identification typically involves the examination of features or marks on either bullets or cartridge cases. A piece of unfired ammunition contains four components: (1) a cartridge case, (2) a primer, (3) propellant (gun powder), and (4) a bullet. The cartridge case holds the unit of ammunition together with the bullet in its mouth. When an individual pulls the trigger of a firearm, a firing pin strikes the primer, which is at the head of the cartridge case. Striking the primer creates a spark that ignites the propellant. The ignition of the propellant forces the bullet to detach from the cartridge case and exit the barrel of the firearm. All of these operations have the potential to impart marks on the cartridge case, on the bullet, or on both. For example, manufacturers use firing pins with different shapes, which are often readily apparent on a fired cartridge case. Similarly, the barrel of the gun has grooves machined into it to impart a spiral spin on the bullet (akin to a football spiral)—some manufactures have different numbers and directions of grooves.

Practitioners call these types of features “class characteristics.”61The official definition used by the professional Association of Firearms and Tool Mark Examiners is “[m]easurable features of a specimen which indicate a restricted group source. They result from design factors and are determined prior to manufacture.” Glossary of the Association of Firearm & Tool Mark Examiners 38 (6th ed. 2013). Class characteristics are the result of design features selected by the manufacturer. For example, a manufacturer may choose to use an elliptical-shaped firing pin or a barrel with six right-hand twisting grooves. The ammunition’s size is also a class characteristic. Class characteristics are a useful first step in firearm examination since observing differences in class characteristics can immediately rule out the possibility that two bullets or cartridge cases were fired by the same gun.

Agreement in class characteristics alone, however, is not sufficient to determine that bullets or cartridge cases were fired by the same gun. To draw that inference, examiners must identify and evaluate “individual characteristics,” which are defined by the AFTE as:

Marks produced by the random imperfections or irregularities of tool surfaces. These random imperfections or irregularities are produced incidental to manufacture and/or caused by use, corrosion, or damage. They are unique to that tool to the practical exclusion of all other tools.62Id. at 65.

Examiners rely on training and experience to assess whether striations are uniquely the result of a particular firearm (in other words, individual characteristics), as opposed to incidental striations that occurred during production and may be apparent in many different firearms of the same class.63These incidental striations are often called “subclass characteristics,” or features that may be produced during manufacture that are consistent among items fabricated by the same tool in the same approximate state of wear. These features are not determined prior to manufacture and are more restrictive than class characteristics. Subclass characteristics can easily be confused with individual characteristics. See Gene C. Rivera, Subclass Characteristics in Smith & Wesson SW40VE Sigma Pistols, 39 AFTE J. 247 (2007). Examiners following the AFTE protocol can reach one of several conclusions based on their evaluation of the individual characteristics: identification, elimination, inconclusive, or unsuitable for comparison.

There are no numeric thresholds for how many individual characteristics must be observed before the examiner can declare that two bullets or cartridge cases were fired by the same gun (that is, “an identification”). Rather, the AFTE protocol states that an identification can be reached “when the unique surface contours of two toolmarks are in ‘sufficient agreement.’ ”64AFTE Theory of Identification as it Relates to Toolmarks, The Ass’n of Firearm and Toolmark Examiners, https://afte.org/about-us/what-is-afte/afte-theory-of-identification [https://perma.cc/C498-FRH2]. As defined by the AFTE:

This “sufficient agreement” is related to the significant duplication of random toolmarks as evidenced by the correspondence of a pattern or combination of patterns of surface contours. . . . The statement that “sufficient agreement” exists between two toolmarks means that the agreement of individual characteristics is of a quantity and quality that the likelihood another tool could have made the mark is so remote as to be considered a practical impossibility.65Id.

This criterion of “sufficient agreement” has been roundly criticized by numerous commentators and courts for being “circular.”66See, e.g., PCAST Report, supra note 43, at 60 (“More importantly, the stated method is circular. It declares that an examiner may state that two toolmarks have a ‘common origin’ when their features are in ‘sufficient agreement.’ It then defines ‘sufficient agreement’ as occurring when the examiner considers it a ‘practical impossibility’ that the toolmarks have different origins.”). It is, however, the criterion adopted by the AFTE and widely used by practicing firearm examiners who conduct casework.67Nicholas Scurich, Brandon L. Garrett & Robert M. Thompson, Surveying Practicing Firearm Examiners, 4 For. Sci. Int’l: Synergy 1, 3 (2022).

B.  The Reception of Firearms Experts in U.S. Courts: 1902–1930

While there is increasingly voluminous scholarship regarding the early origins of gun control in the United States, we are not aware of scholarship exploring the early use of experts seeking to link firearms to particular shootings.68Instead, a body of historical work has explored early firearms regulation and related rights. See generally, e.g., Saul Cornell & Nathan DeDino, A Well Regulated Right: The Early American Origins of Gun Control, 73 Fordham L. Rev. 487 (2004); Charles R. McKirdy, Misreading the Past: The Faulty Historical Basis Behind the Supreme Court’s Decision in District of Columbia v. Heller, 45 Cap. U. L. Rev. 107 (2017). In this Section, we detail what we learned from assembling our database of firearms rulings, collected using searches of legal databases and supplemented with unpublished trial court orders where available.69See supra note 28 for a description of the database and a link to it. As we will describe, twenty-nine of the earliest rulings predated Frye v. United States, a 1923 case that formed the basis for the federal standard for judicial review of novel expert evidence: a requirement of “general acceptance” within the relevant scientific community.70Frye v. United States, 293 F. 1013, 1014 (D.C. Cir. 1923). Further, none of the eleven rulings decided from 1923–1930 cited to Frye—we did not see courts relying on the Frye standard until many decades later. Many of these rulings, absent clear rules of evidence concerning expert testimony, instead focused on whether experts could assist or inform the jury.71Today, such a standard is reflected in Federal Rule of Evidence 702(a). See Fed. R. Evid. 702(a) (asking whether “the expert’s scientific, technical, or other specialized knowledge will help the trier of fact to understand the evidence or to determine a fact in issue”). The earliest rulings date back to the 1870s and they were quite mixed on whether it was erroneous or correct to have admitted expert testimony concerning firearms.72The earliest ruling that we located, Moughon v. State, found error to admit the testimony. 57 Ga. 102, 106 (Ga. 1876). So did Brownell v. People, 38 Mich. 732, 738 (Mich. 1878). But see Dean v. Commonwealth, 32 Gratt. 912, 927–28 (Va. 1879) (holding that it was not erroneous to admit firearms comparison testimony); Sullivan v. Commonwealth, 93 Pa. 284, 296–97 (Penn. 1880) (same).

One of earliest reported cases discussing firearms comparison evidence, Commonwealth v. Best,73Commonwealth v. Best, 62 N.E. 748 (Mass. 1902). was written in 1902 by none other than Oliver Wendell Holmes, then the Chief Justice of the Massachusetts Supreme Judicial Court. Best was convicted of murder, and on appeal, argued that certain firearms comparison evidence offered at the trial was erroneous.74Id. at 749–50. The State argued at trial that Best shot a milkman twice with a Winchester rifle found in Best’s kitchen.75Id. at 750. To prove this, the State fired a third bullet through the gun, took a photograph of it, and published photographs of this bullet and the bullets found in the victim’s body as evidence.76Id.

In conjunction with these photographs, the State called an expert witness to “testif[y] that [the bullets] were marked by rust in the same way that they would have been if they had been fired through the rifle at the farm, and that it took at least several months for the rust that he saw in the rifle to form.”77Id. In other words, the bullets found at the crime scene were rusted only because they were fired through the rusty barrel of Best’s rifle.78Id. Best’s counsel argued—at trial and on appeal—that the evidence was inadmissible because “the conditions of the experiment did not correspond accurately with those of the date of the shooting,” that “the force impelling the different bullets were different in kind,” that “the rifle barrel might be supposed to have rusted more in the little more than a fortnight that had intervened, and that it was fired three times on [the murder date], which would have increased the leading of the barrel.”79Id. To wit: environmental factors called the expert’s conclusion into question.

In his quintessentially succinct style, Justice Holmes swiftly disposed of these arguments, concluding that expert testimony was the only way “the jury could have learned so intelligently how that gun barrel would have marked a lead bullet fired through it,” and “the sources of error suggested were trifling.”80Id. Indeed, despite this being one of the first published opinions that we could find on the admissibility of firearms toolmark evidence, Justice Holmes found “no reason to doubt that the testimony was properly admitted.”81Id. Rejecting the other arguments that Best made on appeal, the court upheld the conviction.82Id.

On the West Coast, two years later, the California Supreme Court decided People v. Weber, a 1906 case that also involved crude firearms comparison evidence. Four members of the Weber family had been killed on their property, three from gunshots and one from blunt force trauma.83People v. Weber, 86 P. 671, 673 (Cal. 1906). Police found a .32-caliber revolver in the basement of the Weber barn with dried blood on it, along with five discarded cartridges.84Id. at 673–74. The defendant was tried and convicted of one of the murders, and he appealed.85Id. at 674. During the trial, the State called “an expert in small arms” who testified that he “compared the markings on the bullets taken from the bodies with the markings on the bullets which he had fired from the pistol,” concluding that these bullets were all fired from the alleged murder weapon.86Id. at 678. While the trial court initially admitted this testimony, the next day, the court struck it, concluding “the comparison of the . . . bullets . . . is not a matter of expert testimony, but one within the ordinary capacities of the average juror or citizen.”87Id. (emphasis added). Thus, the testimony was excluded, but the bullets were all admitted into evidence for the jury to compare during deliberations. On appeal, the California Supreme Court did not disturb the trial court’s ruling, but it did reject the defense’s argument that admitting the bullets into evidence was erroneous.88Id. The court instead held that admitting the evidence to help the jury identify the murder weapon “was pertinent and important.”89Id.

In the 1920s, courts gradually moved toward considering firearms examiners as expert witnesses. In State v. Clark,90State v. Clark, 196 P. 360 (Or. 1921). the Oregon Supreme Court considered a criminal appeal of a manslaughter conviction. Charles Taylor, a worker in Oregon’s National Cascade Forest Reserve, was part of a group assigned to bridge maintenance. Each worker brought a .30-30 Winchester rifle, hoping to hunt “camp meat.”91Id. at 362. One night, Clark and Taylor began hunting, and each fired an initial shot so they could use the spent cartridges as communication whistles.92Id. at 362–63. Taylor then left to hunt, but he was never seen alive again. The subsequent search party found a shell near Taylor’s body and an empty shell in the barrel of Taylor’s gun.93Id. at 367. According to the court, both shells “bore on the brass part of the primer a peculiar mark evidently caused by a flaw in the breechblock of the gun from which they had been fired.”94Id. This design flaw “caused a very slight, almost microscopic protuberance in the primer of the shell, which enlarged photographs ma[de] very clear to the naked eye.”95Id. Law enforcement fired several shots from Clark’s gun, and the cartridges produced the same mark.96Id. Additionally, Clark’s gun created a “sort of double scratch” on the inside of the rim of each shell fired, while Taylor’s gun “made only a single scratch.”97Id. Because of this, the court eliminated “the theory that deceased might have been accidentally shot with his own gun.”98Id. The court held these tests had produced “strong evidence that [Clark] was present and fired the shot that killed Taylor.”99Id. This evidence was presented in trial by the sheriff who described the marks but does not appear to have made more specific conclusions.100Id. at 370. Clark’s counsel objected to this testimony and admission of the photographs, but no specific reason for the objection was provided101Id. The only specific objection regarding the shells was that the photographs were impermissibly enlarged, which the court rejected. Id. at 371.—unsurprising because rules surrounding lay and expert witnesses were less formal in this era. The court held that the testimony was proper and the evidence was admissible.102Id. at 370–71.

In a 1922 case, the Alabama Supreme Court explicitly held—unlike in the cases discussed so far—that firearms comparison examiners could testify as expert witnesses.103Pynes v. State, 92 So. 663, 665 (Ala. 1922). Earlier cases had done so, without much discussion. See, e.g., Sullivan v. Commonwealth, 93 Pa. 284, 296–97 (1880). A person was convicted for killing a man and his dog via gunshot.104Pynes, 92 So. at 665. Police had found a revolver near the victim’s body and the revolver had one cartridge in the chamber that had been discharged.105Id. The State called someone “familiar with such things, [who] had used pistols and shells a good deal,” to testify as an expert.106Id. This expert claimed the casing in the empty chamber and the barrel of the revolver demonstrated it “had not been discharged recently.”107Id. The defense unsuccessfully objected, arguing that the person was not an expert.108Id. On appeal, the Alabama Supreme Court upheld the testimony: “A witness may have expert knowledge of some of the more ordinary affairs of life.”109Id. For a case from the next year finding a similar expert “competent” and any error harmless, see Laney v. United States, 294 F. 412, 416 (D.C. Cir. 1923).

In a 1923 case, however, the Illinois Supreme Court powerfully objected to expert evidence on firearms comparison.110People v. Berkman, 139 N.E. 91, 94–95 (Ill. 1923). The court reversed the conviction on appeal for multiple reasons,111Id. at 94. but it particularly took issue with the State’s use of a police officer as an expert. At trial, a police officer testified for the State that a gun in evidence was the one fired at the victim because it “was the identical revolver from which the bullet introduced in evidence was fired on the night [the victim] was shot.”112Id. The officer was “asked to examine the Colt automatic .32 aforesaid, and gave it as his opinion that the bullet introduced in evidence was fired from the Colt automatic revolver in evidence.”113Id. The Court also questioned the qualifications of the officer:

The state sought to qualify [the officer] for such remarkable evidence by having him testify that he had had charge of the inspection of firearms for the last 5 years of their department; that he was a small-arms inspector in the National Guard for a period of 9 years; and that he was a sergeant in the service in the field artillery, where the pistol is the only weapon the men have, outside of the large guns or cannon.

Id.
The court emphasized:

He even stated positively that he knew that that bullet came out of the barrel of that revolver, because the rifling marks on the bullet fitted into the rifling of the revolver in question, and that the markings on that particular bullet were peculiar, because they came clear up on the steel of the bullet.114Id. (emphasis added).

The court elaborated:

The evidence of this officer is clearly absurd, besides not being based upon any known rule that would make it admissible. If the real facts were brought out, it would undoubtedly show that all Colt revolvers of the same model and of the same caliber are rifled precisely in the same manner, and the statement that one can know that a certain bullet was fired out of a 32-caliber revolver, when there are hundreds and perhaps thousands of others rifled in precisely the same manner and of precisely the same character, is preposterous.115Id.

Finally, the court focused on lay versus expert opinions:

Mere opportunity does not change an ordinary observer into an expert, and special skill does not entitle a witness to give an opinion, when the subject is one where the opinion of an ordinary observer is admissible, or where the jury are capable of forming their own conclusions from the pertinent facts susceptible of proof in common form. . . . If any facts pertaining to the gun and its rifling existed by which such fact could be known, it would have been proper for the witness to have stated such facts and let the jury draw their own conclusions.116Id. at 95 (emphasis added).

The court thus strongly rejected admitting an expert to opine on such firearms evidence.117Id.

By the late 1920s, however, judicial rulings began to shift as the work of Major Goddard became more known. Goddard founded a private crime laboratory—“The Bureau of Forensic Ballistics”118For a detailed account, see Heather Wolffram, Teaching Forensic Science to the American Police and Public: The Scientific Crime Detection Laboratory, 1929-1938, 11 Acad. Forensic Path 52, 55 (2021).—and published the American Journal of Police Science. Goddard became particularly well-known for assisting with the investigation in the Sacco and Vanzetti case in Massachusetts and in the St. Valentine’s Day Massacre in Chicago in 1929.119Id. Before Goddard published his seminal article on ballistic evidence for the U.S. Army in 1925, Forensic Ballistics, many judges, as described above, viewed firearms comparison as a crude technique that jurors could conduct themselves by visually examining the evidence.120Id.

This began to change. For example, in a 1928 Kentucky case, Jack v. Commonwealth, the state supreme court discussed firearms comparison testimony and found the evidence “important if competent, but highly prejudicial  if incompetent.”121Jack v. Commonwealth, 1 S.W.2d 961, 963 (Ky. 1928). The court discussed an article by Major Goddard in Popular Science Monthly122Citing Goddard’s article, the court stated that “the subject of ballistics . . . has reached the status of an exact science.” Id. at 963. and summarized the process:

[T]here is in use a special microscope consisting of two barrels so arranged that both are brought together in one eyepiece. The fatal bullet is placed under one of these barrels, and a test bullet that has been fired through defendant’s pistol is placed under the other barrel, and this brings the sides of the two bullets together and causes them to fuse into one object. If the grooves and other distinguishing marks on both bullets correspond, it is said to show that both balls were fired from the same pistol.123Id. at 963–64.

The court concluded:

It thus appears that this is a technical subject, and in order to give an expert opinion thereon a witness should have made a special study of the subject and have suitable instruments and equipment to make proper test . . . . Clearly the witnesses in this case were not qualified to give such opinions and conclusions and the admission of such evidence was erroneous and prejudicial.124Id. at 964 (emphasis added).

The court therefore rejected the testimony not because it doubted the method itself but because the proffered experts did not follow proper practices.

One year after Jack, the Kentucky Supreme Court again examined firearms comparison testimony in Evans v. Commonwealth.125Evans v. Commonwealth, 19 S.W.2d 1091 (Ky. 1929). The defendant, Evans, was indicted for murder of the Pineville, Kentucky chief of police, and he was ultimately convicted of manslaughter.126Id. at 1092. Six shots were fired in the murder, and police had dug up a bullet from the ground near the scene.127Id. Evans’s primary argument on appeal was that firearms comparison evidence was improper, so the court addressed it “with some degree of elaboration.”128Id. at 1093. The court referenced Jack and noted that one month after Jack was published, Major Goddard—who wrote the article referenced by the court in Jack—offered to testify.129Id. Goddard was given the defendant’s automatic .45 pistol, seven cartridges taken from this pistol, six cartridges found at the scene of the crime, and the bullet that police had taken from the dirt.130Id. at 1094. Goddard concluded “that he was convinced that the bullet that had been introduced into evidence had been fired through [Evans’s] pistol.”131Id. (emphasis added). To justify this conclusion, Goddard gave a detailed account of how he compared the different bullets by putting “the two bullets under the two microscopes together, [so that] in the center . . . you see a single bullet. . . . [I]f these bullets were fired through the same pistol they will match . . . .”132Id. at 1095. Goddard testified that he “only required one single test to identify the bullet in evidence as having been fired through the Evans pistol.”133Id. (emphasis added).

During Goddard’s cross-examination, the jury was allowed to examine the evidence using the microscope.134Id. at 1096. The defense objected that Goddard’s conclusion was one of fact that the jury should instead determine.135Id. at 1097. The court rejected this argument.136Id. Interestingly, the court concluded that Goddard’s opinion was an ordinary lay opinion, not that of an expert.137Id. The court compared Goddard’s testimony to that of a lay witness, saying that “he could smell gasoline,” even though “the average man would have great difficulty in telling just how coal oil or gasoline smells, though acquainted with their odors.” Id. Cross-examination was thus a sufficient safeguard, and “rigid adherence” to the rules of evidence “would be subversive of the ends for which they were adopted.”138Id. The defense also objected to the jury looking through the microscopes which the court quickly dismissed as without “well-founded reason.”139Id.

These two Kentucky Supreme Court opinions formed the framework for the modern approach to firearms comparison evidence. Jack demonstrates that courts would not always let a specific person testify as a qualified expert on firearms comparison. But Evans shows that the courts were not concerned about the underlying validity of the methodology of firearms comparisons. If the State could produce a witness in the mold of Major Goddard, following the now-respected comparison microscope methodology, then the testimony would routinely be admitted.

C.  A National Body of Firearms Rulings: 1930s to 1960s

Beginning in the 1930s, judges began to further develop case law in other parts of the country, with new experts testifying. We identified forty rulings from 1931–1970, each set out in our database. During this time period, rulings spread nationally, as judges appear powerfully influenced by Evans,140Evans v. Commonwealth, 19 S.W.2d 1091 (Ky. 1929). which became one of the lodestar cases for adoption of firearms comparison evidence. Use of toolmark evidence for firearms comparison began to be called “accepted” and “well-recognized” as a methodology. As time went on, judges simply cited to Evans and other prototypical early cases to admit expert testimony, and discussion of the merits of firearms comparison methods diminished. Further, defendants increasingly did not challenge the evidence but rather focused on the preservation of evidence or the qualifications of the testifying experts. These challenges were almost always unsuccessful.

In 1937, for example, the Florida Supreme Court briefly concluded that a firearms comparison expert was “fully qualified to testify as an expert . . . and to draw a reliable conclusion as to whether or not the bullet found in the body of the deceased was fired from the pistol introduced in evidence.”141Riner v. State, 176 So. 38, 39–40 (Fla. 1937). In another Missouri case, the expert himself explained that he “was not a ballistic expert,” but he still argued he had “much experience in the work of identifying firearms.”142State v. Couch, 111 S.W.2d 147, 149 (Mo. 1937). Despite this concession, the court concluded that “he was an expert in the identification of firearms and bullets by the comparison method by means of a microscope.”143Id. In 1938, an Oklahoma appellate court further explained:

There were few decisions with reference to the introduction of expert testimony to identify the weapon from which a shot was fired until recent years, but the science of ballistics is now recognized as one of the best methods in ferreting out crime that could not otherwise be detected. Expert evidence to identify the weapon from which a shot was fired is generally admitted under the rules covering other forms of expert testimony, and it is the modern tendency of the courts to allow the introduction of such testimony, where the witness’ preparation as shown by experience and training qualifies him to give expert opinion on firearms and ballistics tests.144Macklin v. State, 76 P.2d 1091, 1095 (Okla. Crim. App. 1938) (emphasis added).

By 1940, experts could cite fifteen years of experience in “the firing of different caliber pistols,” which was enough to qualify a person as a firearms comparison expert.145McGuire v. State, 194 So. 815, 816 (Ala. 1940). In a 1941 case in Virginia, an expert from the FBI testified that he had twenty years of experience, “six of which had been devoted to the examination of firearms.”146Ferrell v. Commonwealth, 14 S.E.2d 293, 295 (Va. 1941). The expert testified that the cartridge he examined was fired by the defendant’s shotgun.147Id. at 296. The reviewing court cited to Evans,148Id. at 297. as courts continued to do. For example, in State v. McKeever,149State v. McKeever, 101 S.W.2d 22 (Mo. 1936). the expert testified this was his 191st trial—the court allowed the evidence to be admitted without discussion, simply citing to Evans.150Id. at 29. Increasingly brief opinions found “no error” in introduction of such testimony.151See, e.g., Pilley v. State, 25 So.2d 57, 60 (Ala. 1946) (“In the introduction of this evidence there was no error.”); Kyzer v. State, 33 So.2d 885, 887 (Ala. 1947) (finding no error without explanation). In Collins v. State, 33 So.2d 18, 20 (Ala. 1947), the court overruled objections to the expert testimony, stating: “We have had occasion several times to consider questions of this sort, and the principles of law applicable to the same have been repeated frequently, so that it will not be necessary to do so again . . . .” Yet, in none of those prior opinions did the court actually repeat or state its reasoning.

There were some outliers. For example, a 1948 New Mexico Supreme Court ruling reversed the admissibility of “ballistic expert” testimony, which allegedly matched a specific gun to the bullet that killed the victim.152State v. Martinez, 198 P.2d 256, 257–61 (N.M. 1948). After being qualified, the expert testified about his methodology, calling the firearm’s marks “absolutely identical.”153Id. at 257–58. The court was concerned that the expert had concluded with statements such as: “I will state positively that the evidence bullet (death bullet) was fired out of State’s Exhibit No. 2, this [defendant’s] gun.”154Id. at 260 (emphasis added). The court emphasized that while firearms comparison is “almost, if not an exact science,” and “judicial notice may be taken” of the method, ballistic experts still must, “like . . . experts generally,” only provide “opinion testimony.”155Id. While “[i]t may be true that such witnesses as Colonel Goddard, who testified in Evans v. Commonwealth and other reported cases, are so skilled in the science of forensic ballistics that the chance of error is negligible,” they are the exception.156See id. at 261 (citation omitted). Yet, “[t]he belief of a witness that his skill is so transcendent that an error in judgment is impossible, may itself be false or a mistake, assuming that the science is exact.”157Id.

In a rare 1951 Georgia Supreme Court case, Henderson v. State, the court excluded firearms comparison testimony due to concerns with the specific expert. The defense attorney asked the expert “why he did not measure the distance and depth of the grooves, and the witness explained by giving the reply that the microscope was the highest and best evidence.”158Henderson v. State, 65 S.E.2d 175, 177 (Ga. 1951). The court held that the answer was not a “response to the question propounded,”159Id. that the right to a “thorough and sifting cross-examination” was violated, and that the judgement should be reversed for a new trial.160Id.

In a Maryland case, the defendant also attacked the State’s firearms comparison testimony.161Edwards v. State, 81 A.2d 631, 635 (Md. 1951). The court emphatically rejected this position:

For many years ballistics has been a science of great value in ferreting out crimes that otherwise might not be solved. When a pistol is fired, a pressure is developed within the shell which drives the bullet out of the barrel, and the shell is driven back against the breech of the pistol with similar force. The markings on the hard breech of the pistol are thereby stamped on the soft butt of the shell. Testimony to identify the weapon from which a shot was fired is admissible where it is shown that the witness offering such testimony is qualified by training and experience to give expert opinion on firearms and ammunition.162Id.

The court cited back to Best and Evans to justify this result, despite the faint marks and acknowledgment that the marks could have been explained by a different type of weapon.163See id. at 635–36 (noting that “it was admittedly possible that the bullets could have been fired from a Luger” rather than the defendant’s gun).

In a 1964 Florida case, the court provided the following explanation about the recognition of firearms comparison testimony:

It is now well established that a witness, who qualifies as an expert in the science of ballistics, may identify a gun from which a particular bullet was fired by comparing the markings on that bullet with those on a test bullet fired by the witness through the suspect gun. An expert will be permitted to submit his opinion based on such an experiment conducted by him. The details of the experiment should be described to the jury.164Roberts v. State, 164 So. 2d 817, 820 (Fla. 1964).

Finally, a 1969 Illinois appellate case offers some of the earliest descriptions of class and individual characteristics, the predominant terminology in modern firearms comparison testimony:

When a weapon is received at the laboratory it is classified as to type, caliber, make and model. Each gun has class characteristics common to its particular make and model. In addition, each gun has its own individual characteristics. . . . After the gun is received at the laboratory, if operable, it is fired into a bullet recovery box. The bullet in question is then compared with the test bullet under a comparison microscope.165People v. O’Neal, 254 N.E.2d 559, 561–62 (Ill. App. Ct. 1969) (emphasis added).

During this time, courts routinely rejected challenges to firearms experts’ qualifications.166See, e.g., United States v. Hagelberger, 9 C.M.R. 226, 233–34 (1952). And expert qualifications only increased: by now, some experts testified that they had worked on “approximately three to four thousand cases of ballistics.”167Gipson v. State, 78 So. 2d 293, 297 (Ala. 1955). Judicial review of forensic evidence in the following decades involved significant deference, with trial courts deferring to the expert witnesses, and then the appellate courts deferring to the trial courts. Often, courts focused on the specific examiner’s experience rather than assessing the field’s foundational validity.168This, however, is not universal. For a more recent ruling, see State v. Raynor, 254 A.3d 874, 887–88 (Conn. 2020) (noting that refusing to consider new information as a scientific field evolves “would transform the trial court’s gatekeeping function . . . into one of routine mandatory admission of such evidence, regardless of advances in a particular field and its continued reliability”).

D.  Pre-Daubert Cases

In the 1970s and 1980s, leading up to the Daubert ruling in 1993, courts routinely admitted firearms expert testimony, often without discussion.169See, e.g., Hampton v. People, 465 P.2d 394, 400 (Colo. 1970) (stating there was no abuse of discretion for admitting a firearm comparison expert’s testimony). For perhaps the first case referring to the discipline as a type of toolmark comparison, see United States v. Bowers, 534 F.2d 186, 193 (9th Cir. 1976). We located only twenty-four such rulings, perhaps because unpublished rulings became far more common given the broader acceptance of such expert testimony. While challenges to expert qualifications typically failed—with courts citing to the experience of the examiner—courts generally expected examiners to also possess specialized training and credentials.170See, e.g., State v. Hunt, 193 N.W.2d 858, 867 (Wis. 1972) (stating “the witness had great experience in the field of ballistics”); Acoff v. State, 278 So. 2d 210, 217 (Ala. 1973) (concluding expert testimony of witness with “more than six years” of firearms comparison training was “properly allowed”); People v. McKinnie, 310 N.E.2d 507, 510 (Ill. App. Ct. 1974) (finding examiners’ “considerable practical experience” was sufficient, despite lack of “scientific” training). But see State v. Seebold, 531 P.2d 1130, 1132 (Ariz. 1975) (affirming exclusion of proffered experts at trial in which one admitted “he was not a scientist or a criminalist” and the second was a gunsmith and gun shop owner who “had no formal education in the field of ballistics and had never testified before in this field”); Cooper v. State, 340 So. 2d 91, 93 (Ala. Crim. App. 1976) (“The State, in attempting to establish Charles Wesley Smith as an expert in ballistics, elicited some general information on his background, but failed to establish many specific facts to support his expertise in the field of ballistics.”); Bowden v. State, 610 So. 2d 1256, 1258 (Ala. Crim. App. 1992) (affirming trial court’s exclusion of firearms expert’s testimony because it was not a “clear abuse of . . . discretion”).

Some courts excluded firearms testimony based on other issues.171See, e.g., Johnson v. State, 249 So. 2d 470, 472 (Fla. Dist. Ct. App. 1971) (reversing admission of firearms testimony because the State could not produce the bullet taken from the deceased for examination). In a federal case, the defendant was denied access to an expert to examine the evidence and testimony, which was found particularly problematic given the quality of the evidence itself, as “seventy-five percent of this slug was destroyed and the identification was made on the remaining 25%.”172Barnard v. Henderson, 514 F.2d 744, 746 (5th Cir. 1975). Other cases relied on the Confrontation Clause, including one in which a police officer testified about a report by an examiner who was not present at trial.173Stewart v. Cowan, 528 F.2d 79, 82–83 (6th Cir. 1976). Still other courts considered whether experts sufficiently described their work.174People v. Miller, 334 N.E.2d 421, 429 (Ill. App. Ct. 1975). Other cases found it sufficient to admit testimony finding similar class characteristics, even when there was not enough information to compare any individual characteristics. See, e.g., State v. Bayless, 357 N.E.2d 1035, 1058–59 (Ohio 1976).

In general, experts continued to reach highly aggressive conclusions that were permitted by courts. For example, the expert in a 1981 Wyoming case resolved, “The markings on the bullets from the home of appellant’s brother matched the markings found on the bullet removed from [the defendant], establishing that they had been fired from the same gun.”175McDaniel v. State, 632 P.2d 534, 535 (Wyo. 1981). In a leading Virginia case, an expert testified he was “certain” one of the bullets removed from the victim’s body was fired from the defendant’s pistol, and there was “no margin of error.”176Watkins v. Commonwealth, 331 S.E.2d 422, 434 (Va. 1985). The defendant argued on appeal that this “no margin of error” statement was impermissible.177Id. The court rejected this argument, simply concluding that the statement went toward the weight of the testimony, not its admissibility.178Id.

Pre-Daubert, some defendants did contest whether firearms experts relied on sufficient facts and data. In an exemplar Utah case, the expert testified at a preliminary hearing that a bullet fired from the alleged murder weapon matched a bullet taken from the victim’s body.179State v. Schreuder, 712 P.2d 264, 268 (Utah 1985). But while he gave this conclusion, he was not able to give “an exact description of the striations, nor did he have photographs of them available with him in court.”180Id. The court rejected arguments that the expert did not have sufficient foundation for his conclusion, holding that the testimony was within the expert’s specialized knowledge.181Id. at 268–69.

II.  MODERN SCIENTIFIC ASSESSMENTS AND GROWING JUDICIAL SKEPTICISM OF FIREARMS EVIDENCE

Following the U.S. Supreme Court’s ruling in 1993 in Daubert v. Merrell Dow Pharmaceuticals, Inc., federal courts began to more carefully scrutinize firearms evidence, although exclusion remained rare.182See, e.g., Melcher v. Holland, No. 12-0544, 2014 U.S. Dist. LEXIS 591, at *42–44, 51 (N.D. Cal. Jan. 3, 2014) (finding no ineffective assistance of counsel highlighting it was correct to admit firearms evidence); United States v. Sebbern, No. 10 Cr. 87, 2012 U.S. Dist. LEXIS 170576, at *21–24 (E.D.N.Y. Nov. 29, 2012) (finding hearing unnecessary when other courts had examined reliability of firearms evidence). Daubert led to the revision of Federal Rule of Evidence 702 in 2000 that established new standards to assess the reliability of scientific expert testimony. Many of the defendants’ objections shifted from concerns about the experts’ qualifications to concerns about the reliability of the methodology and conclusions,183See, e.g., Abruquah v. State, No. 2176, 2020 Md. App. LEXIS 53, at *19–25 (Md. Ct. Spec. App. Jan. 17, 2020) (defense objections regarding methodology and expert conclusion language); United States v. Mouzone, 687 F.3d 207, 215–17 (4th Cir. 2012) (defense objections focused on expert allegedly violating limits imposed by judge on conclusion language). and about the use of inadmissible hearsay evidence as a basis for the experts’ conclusions.184See, e.g., United States v. Corey, 207 F.3d 84, 87–92 (1st Cir. 2000); Green v. Warren, No. 12-6148, 2013 U.S. Dist. LEXIS 179765, at *21–22 (D.N.J. Dec. 20, 2013). At the state level, there was not any immediate difference in how courts approached firearms expert testimony post-Daubert; methodology and expert qualifications were more explicitly mentioned, but the overall analysis largely remained the same.185See, e.g., State v. Gainey, 558 S.E.2d 463, 473–74 (N.C. 2002).

In the late 1990s and early 2000s, courts began rejecting expert firearms comparison testimony as unreliable, largely by relying on Daubert. In our database, we include just seven cases from 1993–2000. However, the number of rulings begins to dramatically increase after 2000, with 188 rulings from 2000 to 2022. We turn next to that rich body of modern case law.

Figure 1.  Reported U.S. Firearms Rulings by Decade

 

Figure 1 illustrates this remarkable trend—one can see a fairly steady number of twenty or fewer reported judicial rulings regarding firearms comparison evidence through the 1990s. Yet, beginning in the early 2000s, these rulings began to dramatically increase in number.

The Supreme Court in Daubert revolutionized judicial review of scientific evidence by setting out five factors for courts to consider in evaluating expert testimony: whether the theory or technique relied on (1) can be (and has been) tested, (2) has been subjected to peer review and publication, (3) has a known or potential rate of error, (4) includes the existence and maintenance of standards controlling its operation, and (5) is generally accepted within the relevant scientific community.186Daubert v. Merrell Dow Pharmaceuticals, Inc., 509 U.S. 579, 593–94 (1993). We provide an overview of each factor and how courts generally have reviewed them in the context of firearms comparison testimony.

First, courts generally have not questioned the “testability” of firearms forensics, a “key question” when examining reliability.187Id. at 593. A series of courts have held that the propositions that “firearms leave discernible toolmarks on bullets and cartridge casings fired from them, and that trained examiners can conduct comparisons to determine whether a particular gun has fired particular ammunition . . . can be, and have been, tested.”188United States v. Tibbs, No. 2016 CF1 19431, 2019 D.C. Super. LEXIS 9, at *25 (D.C. Super. Ct. Sept. 5, 2019); see also United States v. Monteiro, 407 F. Supp. 2d 351, 369 (D. Mass. 2006) (“[T]he existence of the requirements of peer review and documentation ensure sufficient testability and reproducibility to ensure that the results of the technique are reliable.”); United States v. Otero, 849 F. Supp. 2d 425, 433 (D.N.J. 2012) (“Though [it] inherently involves the subjectivity of the examiner’s judgment as to matching toolmarks, the AFTE theory is testable on the basis of achieving consistent and accurate results.”); United States v. Romero-Lobato, 379 F. Supp. 3d 1111, 1118 (D. Nev. 2019) (“There is little doubt that the AFTE method of identifying firearms satisfies [the testing requirement].”); United States v. Ashburn, 88 F. Supp. 3d 239, 245 (E.D.N.Y. 2015) (“The AFTE methodology has been repeatedly tested.”).

Second, many courts have determined the AFTE method of toolmark identification has been subject to sufficient peer review and publication, largely through the AFTE Journal.189See, e.g., Ashburn, 88 F. Supp. 3d at 245–46 (finding AFTE method has been subjected to peer review through the AFTE Journal); Otero, 849 F. Supp. 2d at 433 (describing the Journal’s peer reviewing process and finding the methodology subject to peer review); United States v. Taylor, 663 F. Supp. 2d 1170, 1176 (D.N.M. 2009) (finding AFTE method subjected to peer review through AFTE Journal and two articles submitted by the government in peer-reviewed journal about the methodology); Monteiro, 407 F. Supp. 2d at 366–67 (describing AFTE Journal’s peer reviewing process and finding it meets peer review element). However, courts are beginning to more rigorously inspect the validity of the peer review process at that journal. Prior to January 2020, the AFTE Journal used a highly unusual “open-review” process whereby the identities of the authors and the reviewers were disclosed and direct communication was encouraged. Furthermore, all of the reviewers were members of AFTE who “ha[d] a vested, career-based interest in publishing studies that validate their own field and methodologies.”190Tibbs, 2019 D.C. Super. LEXIS 9, at *33. These factors led a D.C. Superior Court judge to conclude in 2019: “[T]he vast majority of [firearms comparison] studies are published in a journal that uses a flawed and suspect review process, [which] greatly reduces its value as a scientific publication.”191Id. at *35. Therefore, the peer review factor “on its own does not, despite the sheer number of studies conducted and published, work strongly in favor of admission of firearms and toolmark identification testimony.”192Id. at *36. Nevertheless, courts have cited to other studies or reports to validate the soundness of toolmark comparison—one federal court curiously cited to the 2009 NAS and 2016 PCAST reports as evidence of peer review, despite how damning those reviews are of the method.193See Romero-Lobato, 379 F. Supp. 3d at 1119 (D. Nev. 2019) (“[O]f course, the NAS and PCAST Reports themselves constitute peer review despite the unfavorable view the two reports have of the AFTE method. The peer review and publication factor therefore weighs in favor of admissibility.”). But see United States v. Tibbs, No. 2016 CF1 19431, 2019 D.C. Super. LEXIS 9, at *29 (D.C. Super. Ct. Sept. 5, 2019) (“If negative post-publication commentary from an external reviewing body can satisfy this prong of the Daubert analysis, then the peer reviewed publication component would be more or less read out of Daubert, leaving behind only the requirement of some type of publication.”).

Third, courts have tended to view the error rate for forensics firearms testing as low, though they also sometimes acknowledge that the error rate is “presently unknown.”194United States v. Johnson, No. (S5) 16 Cr. 281 (PGG), 2019 U.S. Dist. LEXIS 39590, at *55 (S.D.N.Y. Mar. 11, 2019) (citing Ashburn, 88 F. Supp. 3d at 246; United States v. Diaz, No. CR 05-00167 WHA, 2007 U.S. Dist. LEXIS 13152, at *27 (N.D. Cal. Feb. 12, 2007)). One federal court concluded that “it is not possible” to calculate an absolute error rate for firearms analysis because “the process is so subjective and qualitative.”195United States v. Monteiro, 407 F. Supp. 2d 351, 367 (D. Mass. 2006). This third factor is particularly important for rigorous assessment because “an expert witness’s ability to explain the methodology’s error rate—in other words, to describe the limitations of her conclusion—is essential to the jury’s ability to appropriately weigh the probative value of such testimony.”196Tibbs, 2019 D.C. Super. LEXIS 9 at *37. Faced with numerous studies purporting extremely low error rates, many courts have simply accepted the validity of these conclusions that forensics firearms testing does have a nominal error rate197See Ashburn, 88 F. Supp. 3d at 246 (“[T]he error rate, to the extent it can be measured, appears to be low, weighing in favor of admission.”); United States v. Otero, 849 F. Supp. 2d 425, 433–34 (D.N.J. 2012) (summarizing several studies indicating a low error rate); United States v. Taylor, 663 F. Supp. 2d 1170, 1177 (D.N.M. 2009) (“[T]his number [less than 1%] suggests that the error rate is quite low.”); Monteiro, 407 F. Supp. 2d at 367–68 (summarizing relevant studies and finding that the known error rate is not “unacceptably high”). or has “a false positive rate of 1.52%.”198Romero-Lobato, 379 F. Supp. 3d at 1120.

In more recent years, as we discuss in a later Section in more detail, courts have begun to reexamine the validity of the error studies and rates presented.199See infra Section II.D; State v. Terrell, No. CR170179563, 2019 Conn. Super. LEXIS 827, at *3 (Conn. Super. Ct. Mar. 21, 2019) (“[The toolmark field] is also not static. A methodology may at one time be viewed as reliable by the scientific community and later fall out of favor.”). Citing basic design flaws of most studies in the field and the studies’ failure to address a large number of “inconclusive” results, one court, for example, found “it difficult to conclude that the existing studies provide a sufficient basis to accept the low error rates for the discipline that these studies purport to establish.”200Tibbs, 2019 D.C. Super. LEXIS 9 at *40–41. Other courts noted concerns with the lack of rigorous testing but did not find this sufficiently persuasive to exclude the evidence outright.201Romero-Lobato, 379 F. Supp. 3d at 1120 (“While the Court is cognizant of the PCAST Report’s repeated criticisms regarding the lack of true black box tests, the Court declines to adopt such a strict requirement for which studies are proper and which are not. Daubert does not mandate such a prerequisite for a technique to satisfy its error rate element.”).

Fourth, many judges have focused on how the AFTE methodology lacks clearly defined, objective standards. Judges have variously described the AFTE method as “inherently vague,”202United States v. Glynn, 578 F. Supp. 2d 567, 572 (S.D.N.Y. 2002). “more of a description of the process of firearm identification rather than a strictly followed charter for the field,”203United States v. Monteiro, 407 F. Supp. 2d 351, 371 (D. Mass. 2006). and “merely unconstrained subjectivity masquerading as objectivity.”204Tibbs, 2019 D.C. Super. LEXIS 9 at *69. And as many courts have pointed out, “the AFTE standard is circular—an identification can be made upon sufficient agreement, and agreement is sufficient when an identification can be made.”205People v. Ross, 129 N.Y.S. 3d 629, 634 (N.Y. Sup. Ct. 2020); see also United States v. Taylor, 663 F. Supp. 2d 1170, 1177 (D.N.M. 2009) (“[T]he AFTE theory is circular.”); Monteiro, 407 F. Supp.2d at 370 (“[T]he AFTE Theory . . . is tautological.”); United States v. Green, 405 F. Supp. 2d 104, 114 (D. Mass. 2005) (stating the method is “either tautological or wholly subjective”). The inherent subjectivity has weighed against admissibility of firearms comparison evidence for many courts.206See, e.g., Romero-Lobato, 379 F. Supp. 3d at 1121 (“With the AFTE method, matching two tool marks essentially comes down to the examiner’s subjective judgment based on his training, experience, and knowledge of firearms. This factor weighs against admissibility.”); United States v. Ashburn, 88 F. Supp. 3d 239, 246–47 (E.D.N.Y. 2015) (discussing subjectivity); Ross, 129 N.Y.S.3d at 633 (describing testimony that “there is no across-the-board standard as to what is ‘sufficient agreement’ in his field”); United States v. Sebbern, No. 10 Cr. 87(SLT), 2012 U.S. Dist. LEXIS 170576, at *11 (E.D.N.Y. Nov. 30, 2012) (“[T]he standards employed by examiners invite subjectivity.”). Courts, however, have often also noted that they find such subjectivity “not fatal” to admissibility.207See Ashburn, 88 F. Supp. 3d at 246–47 (“[T]he subjectivity of a methodology is not fatal under Rule 702 and Daubert.”); Cohen v. Trump, 2016 U.S. Dist. LEXIS 117059, at *35 (S.D. Cal. Aug. 29, 2016) (“[S]ubjective opinions based on an expert’s experience in the industry [are] proper”); Romero-Lobato, 379 F. Supp. 3d at 1120 (“Federal Rule of Evidence 702 inherently allows for an expert with sufficient knowledge, experience, or training to testify about a particular subject matter.”). Thus, courts often note that subjectivity alone does not make a method unreliable and they are focused on evaluating reliability.208See, e.g., Romero-Lobato, 379 F. Supp. 3d at 1120 (“The mere fact that an expert’s opinion is derived from subjective methodology does not render it unreliable.”); United States v. Otero, 849 F. Supp. 2d 425, 431 (D.N.J. 2012) (“[E]xpert testimony on matters of a technical nature or related to specialized knowledge, albeit not scientific, can be admissible under Rule 702, so long as the testimony satisfies the Court’s test of reliability and the requirement of relevance.”).

Finally, the last Daubert factor hinges on general acceptance within the relevant scientific community. Who constitutes the “relevant” scientific community has never been defined with precision, yet it is often determinative. Because the AFTE method is accepted within the organization’s own community of firearms examiners, courts frequently find the requisite general acceptance.209See, e.g., United States v. Shipp, 422 F. Supp. 3d 762, 782 (E.D.N.Y. 2019) (“Most courts have, in cursory fashion, identified toolmark examiners as the relevant community, and have summarily determined that the AFTE Theory is generally accepted in that community.”). But other judges have pointed out that this narrow definition is comprised exclusively of individuals “whose professional standing and financial livelihoods depend on the challenged discipline.”210United States v. Tibbs, No. 2016 CF1 19431, 2019 D.C. Super. LEXIS 9, at *73 (D.C. Super. Ct. Sept. 5, 2019); see also Shipp, 422 F. Supp. 3d at 783 (“The AFTE Theory has not achieved general acceptance in the relevant community.”). One court notes, “It is self evident that practitioners accept the validity of the method as they are the ones using it. Were the relevant scientific community limited to practitioners, every scientific methodology would be deemed to have gained general acceptance.” State v. Terrell, No. CR170179563, 2019 Conn. Super. LEXIS 827, at *14 (Conn. Super. Ct. Mar. 21, 2019). In other forensics fields, acceptance among only practitioners has been deemed unreliable and has led to the exclusion of the evidence under Daubert. See, e.g., United States v. Saelee, 162 F. Supp. 2d 1097, 1104 (D. Alaska 2001) (“[G]eneral acceptance of the theories and techniques involved in the field . . . among the closed universe . . . proves nothing.”). Thus, perhaps the relevant scientific community should be broadened to include nonpractitioner research scientists.

While acknowledging the discipline’s weaknesses, most federal courts have balanced the Daubert factors and found testimony admissible. As one federal court put it: “[T]his lack of objective criteria is countered by the method’s relatively low rate of error, widespread acceptance in the scientific community, testability, and frequent publication in scientific journals.”211Romero-Lobato, 379 F. Supp. 3d at 1122; see also Ricks v. Pauch, No. 17-12784, 2020 U.S. Dist. LEXIS 50109 (E.D. Mich. Mar. 23, 2020) (“Given that no court has ever found Firearm and Toolmark Identification evidence to be inadmissible under Daubert, it is clear that firearm identification testimony meets the Daubert reliability standards and can be admitted as evidence.” (quoting United States v. Alls, No. CR2-08-223 (S.D. Ohio Dec. 7, 2009))); United States v. Wrensford, No. 2013-0003, 2014 U.S. Dist. LEXIS 102446, at *57 (D.V.I. July 28, 2014) (finding “consistent with other courts—that the concerns with subjectivity as it may impact testability, standards, and protocols do not tip the scales against admissibility”). Further, as noted, Rule 702 was revised in 2000 to incorporate Daubert, but it specified additional factors, including asking courts to examine the application of a method to the facts in a case.212Fed. R. Evid. 702. Courts vary in whether they simply consider Daubert factors alone,213See, e.g., United States v. Chavez, No. 15-CR-00285-LHK-1, 2021 U.S. Dist. LEXIS 237830, at *17 (N.D. Cal. Dec. 13, 2021) (finding that four of five Daubert factors weighed in favor of admissibility). or whether they also discuss Rule 702—as will be discussed next, litigants have increasingly focused on the as-applied language in Rule 702, critiquing how the method was used, as well as on the language an expert used to express conclusions.

A.  Post-Daubert Cases

As a federal district court noted in 2005, for over a decade after the Daubert ruling, “every single court post-Daubert has admitted [firearms identification] testimony, sometimes without any searching review, much less a hearing.”214United States v. Green, 405 F. Supp. 2d 104, 108 (D. Mass. 2005) (emphasis omitted). When courts did examine firearms evidence, early post-Daubert challenges often focused on if the expert’s qualifications were sufficient under Rule 702,215For a case affirming disqualification of defense, not prosecution expert, see State v. Hurst, 828 So. 2d 1165 (La. Ct. App. 2002). even if they did begin to discuss questions regarding reliability of methods and principles used.216See, e.g., State v. Samonte, 928 P.2d 1, 26–27 (Haw. 1996) (discussing the defendant’s argument that prosecution’s firearms expert was not qualified); Whatley v. State, 509 S.E.2d 45, 50 (Ga. 1998) (rejecting the defendant’s argument that evidence used was “inherently unreliable” and noting the “ballistics evidence introduced in this case is not novel”). But see Sexton v. State, 93 S.W.3d 96, 101 (Tex. Ct. Crim. App. 2002) (rejecting expert’s claim that the technique was “one hundred percent accurate” and noting while the “underlying theory of toolmark examination could be reliable in a given case,” the use in this case on unfired bullets was not sufficiently established). Further cases discussed—and rejected—the question of whether an expert’s conclusions were based on inadmissible hearsay, rather than their own observations and conclusions.217See State v. Montgomery, No. 94CA40, 1996 Ohio App. LEXIS 1361, at *14 (Oh. Ct. App. Mar. 29, 1996) (“While it is true that other colleagues provided [the expert] with information . . . the major part of his opinion was based on his own observations and expertise.”). And many courts, both state and federal, continued to admit the testimony without serious discussion.218See, e.g., State v. Gainey, 558 S.E.2d 463, 473–74 (N.C. 2002) (rejecting challenge to prosecution’s expert because of “extensive knowledge of the subject matter”); United States v. O’Driscoll, No. 4:CR-01-277, 2003 U.S. Dist. LEXIS 3370, at *4–6 (M.D. Pa. Feb. 10, 2003) (briefly rejecting challenge); United States v. Foster, 300 F. Supp. 2d 375, 376–77 (D. Md. 2004) (same). But for a particularly detailed review of application of Daubert factors to firearms comparison evidence, see United States v. Hicks, 389 F.3d 514, 526 (5th Cir. 2004).

In additional cases, judges dismissed objections to firearms experts whose testimony was said to reach “ultimate issues,” with the judges noting that the experts only opined regarding an acceptable “reasonable scientific certainty.”219State v. Riley, 568 N.W.2d 518, 526 (Minn. 1997). Thus, judges have emphasized the flexibility of the Daubert and Kumho Tire220Kumho Tire Co. v. Carmichael, 526 U.S. 137, 152–53 (1999) (setting out the application of Daubert to expert testimony by nonscientists). standards. As a Southern District of New York ruling explained:

The Court has not conducted a survey, but it can only imagine the number of convictions that have been based, in part, on expert testimony regarding the match of a particular bullet to a gun seized from a defendant or his apartment. It is the Court’s view that the Supreme Court’s decisions in Daubert and Kumho Tire, did not call this entire field of expert analysis into question. It is extremely unlikely that a juror would have the same experience and ability to match two or more microscopic images of bullets.221United States v. Santiago, 199 F. Supp. 2d 101, 111–12 (S.D.N.Y. 2002).

B.  Growing Judicial Skepticism

The federal courts took the lead in beginning to scrutinize firearms comparison testimony more closely. Judges began to write opinions with detailed examinations of the underlying methods experts used. Federal courts then imposed partial exclusions regarding either (1) the methods or qualifications of the particular experts or (2) the language the expert was permitted to use to describe the conclusion. In more recent years, state courts, including trial courts, have joined federal courts in asking more detailed questions and limiting uses of firearms expert testimony.

A turning point was the District of Massachusetts ruling in United States v. Green.222United States v. Green, 405 F. Supp. 2d 104 (D. Mass. 2005). Then-judge Gertner described that the firearms expert had planned to testify about individual characteristics which the expert stated could be matched “to the exclusion of every other firearm in the world.”223Id. at 107. At the opinion’s outset, the court stated that this conclusion was “extraordinary.”224Id. The court also gave one of the earliest detailed descriptions of the exactness—or lack thereof—of the toolmark comparison methodology:

In firearm toolmark comparisons, exact matches are rare. The examiner has to exercise his judgment as to which marks are unique to the weapon in question, and which are not.

In fact, shell casings have myriad markings, some of which appear on all casings from the same type of weapon (“class characteristics”) or those manufactured at the same time (“sub-class characteristics”). Others are arguably unique to a given weapon (“individual characteristics”) or are unique to a single firing (“accidental characteristics”).225Id.

Judge Gertner then explained:

The task of telling them apart is not an easy one. Even if the marks on all of the casings are the same, this does not necessarily mean they came from the same gun. Similar marks could reflect class or sub-class characteristics, which would define large numbers of guns manufactured by a given company. Just because the marks on the casings are different does not mean that they came from different guns. Repeated firings from the same weapon, particularly over a long period of time, could produce different marks as a result of wear or simply by accident.226Id.

Judge Gertner emphasized that in “distinguishing class and sub-class characteristics from individual ones,” the examiner “conceded, over and over again, that he relied mainly on his subjective judgment. There were no reference materials of any specificity, no national or even local database on which he relied.”227Id. Despite these concerns, the court candidly acknowledged that “the problem for the defense is that every single court post-Daubert has admitted this testimony, sometimes without any searching review, much less a hearing.”228Id. Judge Gertner ultimately allowed the expert testimony because “any other decision [would] be rejected by appellate courts, in light of precedents across the country.”229Id. at 109. Nevertheless, the court did not “allow [the expert] to conclude that the match he found by dint of the specific methodology he used permits ‘the exclusion of all other guns’ as the source of the shell casings.”230Id. at 124.

In a second Massachusetts case, United States v. Monteiro, Judge Gertner next held—for the first time—that firearms comparison evidence was inadmissible on an as-applied challenge under Rule 702.231United States v. Monteiro, 407 F. Supp. 2d 351, 375 (D. Mass. 2006). Because of “the extensive documentary record,” the court held that the “underlying scientific principle behind firearm identification—that firearms transfer unique toolmarks to spent cartridge cases—is valid under Daubert.”232Id. at 355. At the same time, Judge Gertner noted that the “process of deciding that a cartridge case was fired by a particular gun is based primarily on a visual inspection” that is “largely a subjective determination.”233Id. (emphasis added). Because of this subjectivity, a testifying examiner must “follow the established standards for intellectual rigor in the toolmark identification field with respect to documentation of the reasons for concluding there is a match (including, where appropriate, diagrams, photographs or written descriptions), and peer review of the results by another trained examiner in the laboratory.”234Id. Ultimately, the court concluded that even though the methodology could be reliable and even though the examiner was qualified based on his training and experience, the expert’s opinion was inadmissible because the expert did not sufficiently comply with proper peer review and documentation requirements.235Id. The Government, however, was allowed—without prejudice—to resubmit evidence of the test results that complied with the standards in the field. Id.

Other federal courts began to follow the approach of Judge Gertner. The Northern District of California in 2007 held that an expert could only testify to a “reasonable degree of certainty in the ballistics field.”236United States v. Diaz, No. CR 05-00167 WHA, 2007 U.S. Dist. LEXIS 13152, at *3 (N.D. Cal. Feb. 12, 2007). But the court commented:

[I]t is important to note that—at least according to this record—there has never been a single documented decision in the United States where an incorrect firearms identification was used to convict a defendant. This is not to say that examiners do not make mistakes. The record demonstrates that examiners make mistakes even on proficiency tests. But, in view of the thousands of criminal defendants who have had an incentive to challenge firearms examiners’ conclusions, it is significant that defendants cite no false-positive identification used against a criminal defendant in any American jurisdiction.237Id. at *41.

Other federal courts, however, instead continued to admit conclusions given with “100% degree[s] of certainty.”238United States v. Natson, 469 F. Supp. 2d 1253, 1261 (M.D. Ga. 2007); see also United States v. Williams, 506 F.3d 151, 161 (2nd Cir. 2007) (discussing United States v. Santiago, 199 F. Supp. 2d 101 (S.D.N.Y. 2002), and agreeing that firearms comparison testimony remains proper). For a state court case discussing Monteiro and emphasizing that California admissibility standards are different, see People v. Gear, No. C049666, 2007 Cal. App. Unpub. LEXIS 6454 (Cal. Ct. App. Aug. 8, 2007). The next shift occurred after the scientific community produced substantial reports raising new reliability questions.

C.  The 2008, 2009, and 2016 Scientific Reports

Over half of the rulings in our database occurred after 2009 when the National Academy of Sciences released a groundbreaking report concerning forensic evidence. To be sure, commercial legal databases may have a greater concentration of more recent appellate rulings. But one might have expected a similar outpouring of judicial rulings after the Daubert ruling in 1993—a fairly modern opinion. Instead, we observe change following an intervention by the scientific community over a decade and a half later.

During this time, the separate field of comparative bullet lead analyses—in which examiners claimed to use chemistry to identify unique elemental makeup of a bullet—was discredited and abandoned by the FBI after the NAS found it lacked any scientific foundation.239Nat’l Rsch. Council, Forensic Analysis: Weighing Bullet Lead Evidence 6 (2004). The NAS is “a private, nonprofit, self-perpetuating society of distinguished scholars engaged in scientific and engineering research, dedicated to the furtherance of science and technology and to their use for the general welfare.”240See 2008 NAS Report, supra note 40, at iii. Indeed, even before the report, courts had begun to exclude such evidence.241See, e.g., Clemons v. State, 896 A.2d 1059, 1074–79 (Md. 2006); Ragland v. Commonwealth, 191 S.W.3d 569, 574–80 (Ky. 2006). While it was a very different discipline, those developments may have raised further concerns in the judiciary regarding the work of firearms examiners.

In a 2008 report focused on the feasibility of a national ballistic imaging database, the NAS concluded that underlying assumptions of firearms comparisons were not yet validated.2422008 NAS Report, supra note 40, at 3 (“The validity of the fundamental assumptions of uniqueness and reproducibility of firearms-related toolmarks has not yet been fully demonstrated.”). Furthermore, “a significant amount of research” would need to be done to determine what characteristics might allow one to determine a probative connection between pieces of firearms evidence.243Id.; see also United States v. Taylor, 663 F. Supp. 2d 1170, 1175 (D.N.M. 2009) (describing the scope of that report which focused on feasibility of a ballistics database but noting that the question “was inextricably intertwined with the question of ‘whether a particular set of toolmarks can be shown to come from one weapon to the exclusion of all others’ ”).

In 2009, the NAS released its landmark report, Strengthening Forensic Science in the United States, after Congress directed NAS to undertake the study recognizing that substantial improvements were needed in the field of forensic science.244See 2009 NAS Report, supra note 41, at xix. The 2009 NAS Report contains a scientific assessment of a variety of forensic science disciplines along with recommendations for improvements in each discipline and to the forensic system as a whole. The Committee assembled by the NAS included prominent forensic scientists, research scientists, lawyers, and judges.245See id. at xix–xx. The Report identified a wide range of methodological issues with the practices of forensic firearm and toolmark identification.

Although the NAS Report did acknowledge that class characteristics are helpful in narrowing the pool of firearms that may have fired a particular bullet or cartridge case, it recognized that firearm examiners necessarily go beyond class characteristics when making an identification. The Report noted that a “fundamental problem with toolmark and firearms analysis is the lack of a precisely defined process”246Id. at 155. and that the AFTE methodology “does not even consider, let alone address, questions regarding variability, reliability, repeatability, or the number of correlations needed to achieve a given degree of confidence.”247Id. The Report concluded that “[b]ecause not enough is known about the variabilities among individual tools and guns, [firearm examiners are] not able to specify how many points of similarity are necessary for a given level of confidence in the result.”248Id. at 154.

Building on this work by the NAS, the 2016 President’s Council of Advisors on Science and Technology (“PCAST”) published its 2016 report on the use of forensic science in criminal proceedings. The report was a response to President Obama’s question “whether there [we]re additional steps on the scientific side, [in addition to those identified in the 2009 NAS Report], that could help ensure the validity of forensic evidence used in the Nation’s legal system.”249PCAST Report, supra note 43, at x. The advisory group consisted of “leading scientists and engineers, appointed by the President to augment the science and technology advice available to him from inside the White House, and from cabinet departments and from other Federal agencies.”250Id. at iv. The group focused on six feature-comparison methods including firearms-comparison evidence.251Id. at 7.

Consulting with forensic scientists, PCAST reviewed more than two thousand studies from various disciplines.252Id. at 2. The field had responded to the NAS reports by conducting new studies, and PCAST undertook a deep examination of them. As the NAS had done in its 2009 Report, PCAST asked whether each discipline met basic requirements for scientific validity, which consists of both “foundational validity”—whether the method can, in principle, be reliable—and “validity as applied”—whether the method has been reliably applied in practice.253Id. at 47–48, 56–58.

To be foundationally valid, a method must have been subject to “empirical testing by multiple groups, under conditions appropriate to its intended use.”254Id. at 5. Specifically, “the procedures that comprise it must be shown, based on empirical studies, to be repeatable, reproducible, and accurate, at levels that have been measured and are appropriate to the intended application.”255Id. at 47. The studies must also provide “valid estimates of the method’s accuracy,” demonstrating how often an examiner is likely to draw the wrong conclusion even when applying the method correctly (that is, a scientifically valid error rate).256Id. at 5. As PCAST explained, “Without appropriate estimates of [the method’s] accuracy, an examiner’s statement that two samples are similar—or even indistinguishable—is scientifically meaningless: it has no probative value, and considerable potential for prejudicial impact.”257Id. at 6.

Ultimately, as described below, PCAST concluded that all but one of the existing studies did not use appropriate designs to truly test the ability of a firearm examiner to make accurate identifications. PCAST went on to conclude that “[b]ecause there has been only a single appropriately designed study, the current evidence falls short of the scientific criteria for foundational validity.”258Id. at 111. Much like the NAS report that preceded it, PCAST pointed to the necessity for additional, appropriately designed studies to test the validity of firearm examination.259Id.

1.  Evaluation of the Scientific Studies

PCAST divided the firearms identification studies it reviewed into two different types: set-to-set studies and sample-to-sample studies. In a set-to-set study, examiners are given two sets of bullets and then asked to link the first set of bullets to the second set of bullets. In a sample-to-sample study, examiners are given two bullets to compare and are asked to judge whether the bullets were fired by the same gun or not. This process is then repeated for other test sets of bullets. PCAST concluded that “set-based studies are not appropriately-designed black-box studies from which one can obtain proper estimates of accuracy.”260Id. at 106.

The principal problem of set-to-set studies is that test takers can leverage the design to gain inferences about other comparisons, making the task totally unlike real-world comparison work.261United States v. Cloud, 576 F. Supp. 3d 827, 842–43 (E.D. Wash. 2021) (“Such studies lack external validity, as examiners conducting real-world comparisons have neither the luxury of knowing a true match is somewhere in front of them nor of making process-of-elimination-type inferences to reach their conclusions.”). For example, if an examiner identifies a match between bullets one and two, and then determines that bullet one and bullet A match, then bullet two and bullet A must also be a match by implication. Thus, a test taker would get a correct response for linking unknown bullet two to known bullet A despite never directly comparing the bullets. PCAST noted that: “[t]he Director of the Defense Forensic Science Center analogized set-based studies to solving a ‘Sudoku’ puzzle, where initial answers can be used to help fill in subsequent answers.”262PCAST Report, supra note 44, at 106. Because of this, set-to-set studies typically yield errors rates of zero and very few inconclusive responses.263Id.

At the time PCAST conducted its analysis, there was only a single sample-to-sample study available for firearms identification. The unpublished study was conducted by researchers at the Ames Laboratory in Iowa.264David P. Baldwin, Stanley J. Bajic, Max Morris & Daniel Zamzow, A Study of False-Positive and False-Negative Error Rates in Cartridge Case Comparisons (2014) [hereinafter Ames I]. In this first Ames Lab study, 218 firearm examiners were mailed a test packet that contained cartridge cases to examine. Each test packet totaled 15 separate comparisons for the examiners to evaluate. Unbeknown to the participants, 10 of the comparisons were different-source comparisons, for which the correct response was elimination, and 5 were same-source comparisons, for which the correct response was identification. Examiners were instructed to work alone on the test and to follow the AFTE protocol.

The study reported a 1.01% false positive error rate.265Id. at 3. What was not stated explicitly in the study is that 33.7% of the responses were deemed inconclusive—a pattern of results wildly at odds with the results from the set-to-set studies.266Id. at 16. There were 2,180 different source comparisons of which 735 were inconclusive (735 / 2,180 = 33.7%). Compare that figure to a well-known set-to-set study by Hamby which reported only 8 inconclusives—0.1%—out of 7,605 comparisons. J.E. Hamby, David J. Brundage & James W. Thorpe, The Identification of Bullets Fired From 10 Consecutively Rifled 9mm Ruger Pistol Barrels: A Research Project Involving 507 Participants from 20 Countries, 41 AFTE J. 99 (2009). PCAST noted that “the closed-set studies show a dramatically lower rate of inconclusive examinations and of false positives. With this unusual design, examiners succeed in answering all questions and achieve essentially perfect scores. In the more realistic open designs, these rates are much higher.”267PCAST Report, supra note 44, at 110. PCAST was not the first group to point out the shortcomings of set-to-set studies. The Ames study, for example, stated,

Several previous studies have been carried out to examine this and related issues of individualization and durability of marks [1-5], but the design of these previous studies, whether intended to measure error rates or not, did not include truly independent sample sets that would allow the unbiased determination of false-positive or false-negative error rates from the data in those studies.

Ames I, supra note 264, at 4.

One federal district court, in extensively discussing the PCAST report findings, noted, “Based on the above information, the court finds that the potential rate of error for matching ballistics evidence based on the AFTE Theory does not favor a finding of reliability at this time.”268United States v. Shipp, 422 F. Supp. 3d 762, 778–79 (E.D.N.Y. 2019). The court noted, however, that the FBI and the Ames Laboratory were “currently conducting a second black box study on the AFTE Theory.”269Id. at 779. That study was posted online in early 2021 (and subsequently removed from the Internet).270Components of the Ames II study still appear online. See L. Scott Chumbley, Max D. Morris, Stanley J. Bajic, Daniel Zamzow, Erich Smith, Keith Monson & Gene Peters, Accuracy, Repeatability, and Reproducibility of Firearms Comparisons Part I: Accuracy, https://arxiv.org/ftp/arxiv/papers/2108/2108.04030.pdf [https://perma.cc/EJB6-E434].

The FBI/Ames Laboratory study (hereinafter “Ames II”) utilized a design ambitious in size and scope. First, the study contained both cartridge case and bullet comparisons. The vast majority of previous firearms comparison studies examined only cartridge cases. Second, the study consisted of three rounds that attempted to measure accuracy (round one), repeatability (round two), and reproducibility (round three). Repeatability refers to “the ability of an examiner, when confronted with the exact same comparison once again, to reach the same determination as when first examined.”271Stanley J. Bajic, L. Scott Chumbley, Max Morris & Daniel Zamzoe, U.S. Dep’t of Just., Report: Validation Study of the Accuracy, Repeatability, and Reproducibility of Firearm Comparisons, Ames Laboratory 10 (2020) [hereinafter Ames II] (on file with authors). Reproducibility refers to “the ability of a second examiner to evaluate a set previously viewed by a different examiner and reach the same conclusion.”272Id. at 11. No other study had attempted to measure repeatability and reproducibility of firearm examiner judgments.

In round one of the study, 256 active firearm examiners were sent test packets—each test packet contained 15 comparison sets of bullets and 15 comparison sets of cartridge cases. For each comparison, participants were instructed to make a judgement according to the AFTE Range of Conclusions.273The Range of Conclusions includes the following options: (1) Identification, (2a) Inconclusive-A, (2b) Inconclusive-B, (2c) Inconclusive-C, (3) Elimination, and (4) Unsuitable. See Ames I, supra note 264, at 7. Participants were admonished not to discuss their results with anyone else. However, only 173 participants out of 256 returned their test packets. According to the authors, “the overall rate of false positive error rate was estimated as 0.656% and 0.933% for bullets and cartridge cases, respectively, while the rate of false‐negatives was estimated as 2.87% and 1.87% for bullets and cartridge cases, respectively.”274Ames II, supra note 276, at 2. Here again, there was an enormous amount of inconclusive responses: over 50% of the bullet comparisons were deemed inconclusive, and over 42% of the cartridge comparisons were deemed inconclusive.275Id. at 35.

In round two of the study, participants were sent the same test packet they examined previously. Only 105 participants completed this round.276Id. at 39. The percentage of time that examiners reached the same conclusion in round one and round two ranged from 79% to 62%.277Id. at 39. This does not necessarily mean the examiner reached the correct conclusion about two-thirds of the time; rather, it only suggests she reached the same conclusion about two-thirds of the time. According to the authors, a statistical test comparing the “observed agreement” between conclusions reached in round one and in round two to the “expected agreement” “indicat[ed] ‘better than chance’ repeatability.”278Id. at 45. However, two different statisticians concluded that: “[t]he level of repeatability and reproducibility as measured by the between rounds consistency of conclusions would not appear to support the reliability of firearms examination.”279Alan H. Dorfman & Richard Valliant, A Re-analysis of Repeatability and Reproducibility in the Ames-USDOE-FBI Study, 9 Stat. & Pub. Pol’y 175, 178 (2020).

Only 80 participants completed round three of the study.280Ames II, supra note 276, at 15. The percentage of time that 2 different participants examined the same test set and reached the same conclusion ranged from 68% to 31%.281Id. at 47. These latter results are striking. Less than one-third of the time, 2 different participants looked at the same bullets and reached the same conclusion. This means that over two-thirds of the time (69.1%), 2 different participants reached different conclusions when examining the same set of bullets. A statistical test revealed “better than chance” agreement for same-source bullet comparisons but not different-source bullet comparisons.282Id. at 52.

D.  Litigating the Error Rate Studies

The conclusion reached by PCAST that “firearms analysis currently falls short of the criteria for foundational validity”283PCAST Report, supra note 43, at 112. did not go unnoticed by the defense bar. Admissibility challenges to firearm examiner testimony surged—we include more than eighty such cases in our database.284For recent cases in which the defendant challenged firearms testimony, see People v. Ross, 129 N.Y.S.3d 629, 639 (Sup. Ct. 2020); United States v. Tibbs, No. 2016 CF1 19431, 2019 D.C. Super. LEXIS 9 (D.C. Super. Ct. Sept. 5, 2019); United States v. Davis, No. 4:18-cr-00011, 2019 U.S. Dist. LEXIS 155037 (W.D. Va. Sept. 11, 2019); United States v. Shipp, 422 F. Supp. 3d 762 (E.D.N.Y. 2019); United States v. Johnson, No. (S5) 16 Cr. 281 (PGG), 2019 U.S. Dist. LEXIS 39590 (S.D.N.Y. Mar. 11, 2019), aff’d, 861 F. App’x 483 (2d Cir. 2021); United States v. Romero-Lobato, 379 F. Supp. 3d 1111 (D. Nev. 2019); United States v. Shipp, 422 F. Supp. 3d 762 (E.D.N.Y. 2019); State v. Terrell, No. CR170179563, 2019 Conn. Super. LEXIS 827 (Conn. Super. Ct. Mar. 21, 2019); United States v. Simmons, No. 2:16cr130, 2018 U.S. Dist. LEXIS 18606 (E.D. Va Jan. 12, 2018). These challenges often summarized the PCAST analyses and conclusions in arguing that the field failed to pass Daubert’s muster. These challenges, however, almost universally failed. Critics of PCAST sought to characterize the report as authored by outsiders who failed to learn the fundamentals of firearm examination and who committed numerous errors in their own analysis.285For example, the Organization of Scientific Area Committee (“OSAC”) Firearms and Toolmarks Subcommittee issued a formal response in which it claims to catalog “[e]rrors and [o]missions in PCAST [s]ummaries of [f]irearms and [t]oolmarks [v]alidation [s]tudies.” Org. of Sci. Area Comms. (OSAC) Firearms & Toolmarks Subcomm., Response to the President’s Council of Advisors on Science and Technology (PCAST) Call for Additional References Regarding its Report “Forensic Science in the Criminal Courts: Ensuring Scientific Validity of Feature-Comparison Methods” 11 (2016). See also Ass’n of Firearm & Tool Mark Examiners, Response to Seven Questions Related to Forensic Science Posed on November 30, 2015 by The President’s Council of Advisors on Science and Technology (PCAST) (2015). But the tides have recently begun to shift, as courts imposed new, albeit still limited, restrictions on the type of testimony firearm examiners may offer and how they express conclusions.286See, e.g., Tibbs, 2019 D.C. Super. LEXIS 9; Ross, 129 N.Y.S.3d 629. We have identified thirty-seven judicial rulings imposing limitations on firearms comparison testimony and set out each in Appendix A.

Two factors have contributed to the shifting tides. First, in addition to citing the NAS and PCAST reports, attorneys have called mainstream research scientists to testify generally about scientific methods and principles and specifically about the discipline of firearm examination. These experts are not firearm examiners and typically have never conducted a firearm examination.287See generally, e.g., Faigman et al., supra note 45. Much like the practitioner/researcher distinction in medicine, these experts are researchers who study whether the methods employed by the practitioners are effective. These experts are poised to evaluate claims made in court regarding scientific practices.288For example, judges are supposed to consider whether research appears in a “peer-reviewed” scientific journal. See supra notes 189–192 and accompanying text. Most research on firearm examination is published in the AFTE Journal which is touted in court as a “peer-reviewed scientific” journal. See AFTE J., https://afte.org/afte-journal. Upon closer inspection, however, the peer-review process used by the AFTE Journal is highly dissimilar to the usual process that occurs at scientific journals. See Tibbs, 2019 D.C. Super. LEXIS 9, at *25.

The second major factor concerns additional examination of the PCAST-reviewed studies that potentially undermine the reported error rates and the utility of the validation studies. As noted, one-third of the responses in the Ames I study were inconclusive.289See supra note 266 and accompanying text.

What ought to be done with those responses? PCAST ultimately calculated the error rate without considering them. Other firearm studies actually count inconclusive responses as correct responses, based on the logic that “an inconclusive response is not an incorrect response [so they are] totaled with the correct response and figured into the error rate as such.”290Dennis J. Lyons, The Identification of Consecutively Manufactured Extractors, 41 AFTE J. 246, 255 (2009). But what if those responses are errors? The error rate would be as high as 35% in the Ames I study. Other sample-to-sample studies conducted after the PCAST analyses have reported rates of inconclusive responses over 50%.291See Ames II, supra note 271, at 35. Clearly, determining how to count over half of the responses in a validation study is critical.

There are many legitimate reasons to count the inconclusive responses in the Ames I study, including the fact that “[t]he fraction of samples reported as inconclusive cannot be attributed to a large fraction of poorly marked knowns or questioned samples in this group”292Ames I, supra note 264, at 19. and an inconclusive response is also defined by AFTE as an absence of insufficient quality of marking to reach an identification or elimination.293AFTE Range of Conclusions, Ass’n of Firearm & Tool Mark Examiners, https://afte.org/about-us/what-is-afte/afte-range-of-conclusions [https://perma.cc/EJB6-E434] (last visited July 29, 2022). As noted in a 2020 scientific article, a proper study design would include inconclusive test items so that inconclusive responses could be evaluated and incorporated into the error rate.294Itiel E. Dror & Nicholas Scurich, (Mis)use of Scientific Measurements in Forensic Science, Forensic Sci. Int’l: Synergy 333, 335–36 (2020). No study has yet done so and, as a result, error rates observed in the studies span a range so large as to be wholly unhelpful—anywhere from one percent to over fifty percent, depending on whether the responses are dropped or considered as erroneous. Thus, as one district court recently put it,

But providing examiners in the study setting the option to essentially “pass” on a question, when the reality is that there is a correct answer—the casing either was or was not fired from the reference firearm—fundamentally undermines the study’s analysis of the methodology’s foundational validity and that of the error rate.295United States v. Cloud, 576 F. Supp. 3d 827, 843 (E.D. Wash. 2021).

This crucial issue of inconclusive responses was never considered prior to Tibbs, discussed earlier in this Article,296See United States v. Tibbs, No. 2016 CF1 19431, 2019 D.C. Super. LEXIS 9, at *56–66 (D.C. Super. Ct. Sept. 5, 2019) (discussing the issue of inconclusiveness in an order following an admissibility hearing). in which a defense expert raised the concern during an admissibility hearing. The judge in Tibbs called it “perhaps [the] most substantial issue related to the studies proffered to support the reliability of firearms and toolmark analysis”297Id. at 56–57. and noted that “the methods used in the proffered laboratory studies make a compelling case that inconclusive should not be accepted as a correct answer in these studies.”298Id. at 57–58. To be sure, in one 2020 Washington, D.C. case, a judge discounted those findings for which no defense expert was presented to explain these error rate issues.299See United States v. Harris, 502 F. Supp. 3d 28, 35 (D.D.C. 2020).

Then again, another 2020 case in Oregon limited the admissibility of firearms testimony without the benefit of a defense expert witness.300United States v. Adams, 444 F. Supp. 3d 1248 (D. Or. 2020). This judge expressed major concerns about inconclusive responses in firearms comparison studies and their impact on reported error rates:

It appears to be the case that the only way to do poorly on a test of the AFTE method is to record a false positive. There seems to be no real negative consequence for reaching an answer of inconclusive. Since the test takers know this, and know they are being tested, it at least incentivizes a rate of false positives that is lower than real world results. This may mean the error rate is lower from testing than in real world examinations.301Id. at 1265.

A litany of other concerns besides the inconclusive response issue have been raised about the error rate studies. We mention four important issues here.

First, and most fundamentally, none of the studies were test-blind—the participants knew that they were being tested. There is powerful evidence that human subjects are predictably biased—and behave differently—when they know that they are being tested. The PCAST report emphasized the need for blind testing of forensic techniques.302PCAST Report, supra note 43, at 58–59. So have a host of researchers based on a large body of research documenting the manner in which cognitive biases can lead forensic examiners to make errors.303See generally, e.g., Itiel E. Dror, Cognitive and Human Factors in Expert Decision Making: Six Fallacies and the Eight Sources of Bias, 92 Analytical Chemistry 7998 (2020). Although blind testing is standard in medicine, it has never been standard in error rate studies in forensics.

Second, many of the volunteer participants in both of the Ames studies simply dropped out or participated but did not complete the test. In the Ames II study, “32% of the 256 examiners receiving their first packets failed to report any results, and another 32% of the 256 dropped out before completing all six mailings.”304Alan H. Dorfman & Richard Valliant, Inconclusives, Errors, and Error Rates in Forensic Firearms Analysis: Three Statistical Perspectives, 5 Forensic Sci. Int’l: Synergy 1, 5 (2022). No analysis of the participants who initiated the study but declined to complete it was conducted.305Id. Attrition bias due to nonrandom dropout is a serious concern that has an unknown impact on the reported error rates. Although one court has noted that the “use of volunteers . . . does not provide the clearest indication of the accuracy of the conclusions that would be reached by average toolmark examiners,”306United States v. Tibbs, No. 2016 CF1 19431, 2019 D.C. Super. LEXIS 9, at *47–48 (D.C. Super. Ct. Sept. 5, 2019). courts have not focused on issues related to participant dropout.

Third, there are also questions about whether the materials being used in the studies, such as the types of firearms and the quality of the fired items, are sufficiently representative to draw inferences about the field writ large. By design, studies should be of varying degrees of difficulty, but unfortunately, “[w]ith a few exceptions, each of the forensic firearms studies to date focuses on a single firearm,” and the exceptions are telling, whereas studies that use different types of firearms have resulted in very different error rates for each type.307See Dorfman & Valliant, supra note 304, at 5 (“The few studies that have carried out comparisons over a variety of guns have displayed marked differences in the ease of coming to correct conclusions.”). Further, if in a study, “an examiner is over and over comparing bullets or cartridge cases from the same brand and model, then he or she can be expected to be picking up nuances along the way. A later comparison will have an advantage over the first. We can expect this to lead to a reduction in sample error rates.”308Id. Unlike other forensic identification fields, none of these studies have used technology or databases to ensure the test items are challenging.309Nicholas Scurich, Inconclusives in Firearm Error Rate Studies are Not “a Pass,” L. Probability & Risk (2022) (“[R]esearchers should intentionally select challenging test items, in a manner similar to Professor Koehler’s exemplary fingerprint examiner study involving ‘close non-matches.’ ”). Nor has there been any careful analysis of how representative or challenging these studies are, and this basic problem has not received the judicial attention that it should.

Finally, judges have not focused on the appalling levels of nonrepeatability and nonreproducibility of firearms work in the Ames II study: “[E]xaminers examining the same material twice, disagree[d] with themselves between 20% and 40% of the time.”310Ames II, supra note 271, at 39 tbl.XI; Dorfman & Valliant, supra note 304, at 6. They disagreed with other examiners even more, up to 69% of the time for nonmatching bullets and up to 60% of the time for nonmatching cartridges.311Dorfman & Valliant, supra note 304, at 6. Although there is spirited debate about inconclusive results and whether or not they constitute errors in a study, these rates of intra- and interparticipant consistency should eclipse that entire discourse—they set a limit on validity and cannot be dismissed as a disagreement about the interpretation of inconclusive responses. Yet, likely because Daubert explicitly mentions error rates, not rates of consistency, courts have yet to grapple with these findings and how they can be reconciled with professed error rates of one percent or less.

All of this said, it is not uncommon for judges to respond to these studies, the PCAST report, and critiques from research scientists dismissively. Judges have commonly relied on precedent to make conflated arguments against the invalidity of the studies. For example, one judge in New York state—a Frye jurisdiction where the standard for expert evidence admissibility is the “general acceptance” of the method within the relevant scientific community—recently emphasized that the acceptance of firearms comparison methods within the community of practitioners is “nearly universal”312State v. Vasquez, No. 2203/2019, at 3 (N.Y. Sup. Ct., July 24, 2022). According to this judge, the relevant scientific community is not “experts in ‘scientific methodology,’ which is to say, scientists,” id. at 2, but rather “trained and accredited experts in the field of microscopic ballistics and forensic firearm and toolmark examination” as well as “non-firearm practitioners enumerated in the multiple validation studies that have been conducted to demonstrate the reliability of the discipline and its examination results,” id. at 3 (emphasis added). Conducting a study to demonstrate a result is not good science. and that “the Appellate Division . . . has repeatedly upheld the admission of ballistics expert testimony without the need for a Frye hearing.”313Id. at 5. But the judge then went on to hold that “the PCAST report has been thoroughly discredited”314Id. at 4. and the “the very type of study called for by PCAST—a ‘black box study’—has, since the time of the PCAST report, been repeatedly utilized to validate firearm and toolmark comparison methodology.”315Id. at 5. There was no engagement with the results of those studies or their limitations. Unfortunately, it is common for judges to rely on precedent as a form of “general acceptance” by the courts and not carefully examine the reliability of scientific evidence.316Stephanie L. Damon-Moore, Trial Judges and the Forensic Science Problem, 92 N.Y.U. L. Rev. 1532, 1564 (2017) (“Ironically, the ultimate safeguard against judicial error—appellate review—may actually discourage judges from gatekeeping effectively.”).

E.  Testimonial Limitations and Post-NAS and PCAST Rulings

In recent years, courts have more rigorously evaluated the field of firearms examination, in contrast to over fifty years in which claims made by firearm examiners regarding the foundational validity were uncritically accepted.317See, e.g., United States v. Shipp, 422 F. Supp. 3d 762, 775 (E.D.N.Y. 2019) (“Even though prior decisions have found toolmark analysis to be reliable, it is incumbent upon this court to thoroughly review the critiques of the AFTE Theory found in the NRC and PCAST Reports.”); United States v. Adams, 444 F. Supp. 3d 1248, 1266 (D. Or. 2020) (concluding that it could not “find that the AFTE method enjoys ‘general acceptance’ in the scientific community”); People v. Ross, 129 N.Y.S.3d 629, 641 (N.Y. Sup. Ct. 2020) (“[B]eyond comparing class characteristics forensic toolmark practice lacks adequate scientific underpinning and the confidence of the scientific community as whole.”). These more searching evaluations have led judges to note limitations and knowledge gaps that had rarely been discussed in judicial opinions. Despite increasing awareness of the limitations of the field, almost all courts have nevertheless found it admissible.318Ricks v. Pauch, No. 17-12784, 2020 U.S. Dist. LEXIS 89453, at *29–32 (E.D. Mich. Mar. 23, 2020); see also United States v. Romero-Lobato, 379 F. Supp. 3d 1111, 1117 (D. Nev. 2019) (“[N]o federal court (at least to the Court’s knowledge) has found the AFTE method to be unreliable under Daubert.”); United States v. Davis, No. 4:18-cr-00011, 2019 U.S. Dist. LEXIS 155037, at *12–15 (W.D. Va. Sept. 11, 2019) (“[N]o federal court has outright barred testimony from a qualified firearm or toolmark identification expert.”). This created a new conundrum for courts: how to admit firearms identification evidence in a way that does not overstate its value or cause the fact finder to be misled. In the Sections that follow, we report the four tacks that courts have taken when admitting firearm examination evidence: (1) limiting the language that experts can use when testifying to their conclusions, (2) limiting conclusions to class characteristics only, (3) ruling that evidence concerning the proficiency of firearms experts is relevant to the preliminary question whether to qualify the expert, and (4) examining the as-applied question whether the method was reliably used in the particular case.

1.  Limiting Conclusion Testimony

While many courts have continued to admit firearms examiner testimony, “[m]any of these courts admitted the proffered testimony only under limiting instruction restricting the degree of certainty to which firearm and toolmark identification specialists may express their identifications.”319Davis, 2019 U.S. Dist. LEXIS 155037, at *15. The case law that has resulted is diverse, sometimes inconsistent, and reflects a gradual evolution of judicial approaches. As we will describe, in general, a range of courts have limited testimony based on the concerns about toolmark identification methodology.320See, e.g., Shipp, 422 F. Supp. 3d at 783 (preventing a toolmark expert from testifying “to any degree of certainty, that the recovered firearm is the source of the recovered bullet fragment or the recovered shell casing”); Adams, 444 F. Supp. 3d at 1266–67 (same); United States v. Monteiro, 407 F. Supp. 2d 351, 373 (D. Mass. 2006) (same); Davis, 2019 U.S. Dist. LEXIS 155037, at *24 (“[W]itnesses may not testify as to a ‘match,’ that the cartridges bear the same ‘signature,’ that they were fired by the same gun, or words to that effect.”); United States v. Glynn, 578 F. Supp. 2d 567, 575 (S.D.N.Y. 2008) (limiting testimony to “be stated in terms of ‘more likely than not,’ but nothing more”).

The earlier decisions had held that an examiner could only testify to a milder degree, forbidding aggressive statements of a match, “the exclusion of all other firearms in the world,”321United States v. Cazares, 788 F.3d 956, 989 (9th Cir. 2015); United States v. Taylor, 663 F. Supp. 2d 1170, 1180 (D.N.M. 2009); United States v. Ashburn, 88 F. Supp. 3d 239, 249 (E.D.N.Y. 2015); see also United States v. Love, No. 2:09-cr-20317-JPM, at 14–15 (W.D. Tenn. Feb. 8, 2011) (excluding testimony with conclusions of absolute or practical certainty). and instead imposing a more cautious formulation, such as a “reasonable degree of ballistic certainty.”322United States v. Diaz, No. CR 05-00167 WHA, 2007 U.S. Dist. LEXIS 13152, at *36 (N.D. Cal. Feb. 12, 2007). Other courts have taken a different approach, using more familiar standards of proof as a frame of reference—courts have ruled that the examiner can only opine that it is “more likely than not” that the bullet recovered from the crime scene came from the defendant’s firearm.”323See Glynn, 578 F. Supp. 2d at 574–75 (limiting testimony to “more likely than not” conclusion). The table below summarizes some of the main approaches that courts have taken toward limiting such testimonial conclusions. Appendix A summarizes all thirty-seven opinions that we have located, through 2022, including unpublished trial court rulings.

Table 1.  Testimonial Limitations on Firearms Examiners
Court-ordered Conclusion LanguageCitations from selected examples
“more likely than not”United States v. Glynn, 578 F. Supp. 2d 567 (S.D.N.Y. 2008)
“reasonable degree of ballistic certainty”United States v. Monteiro, 407 F. Supp. 2d 351 (D. Mass. 2006)
“consistent with”United States v. Sutton, No. 2018 CF1 009709 (D.C. Super. Ct. May 9, 2022)
“a complete restriction on the characterization of certainty”United States v. Willock, 696 F. Supp. 2d 536 (D. Md. 2010)
“the recovered firearm cannot be excluded as the source of the cartridge casing found on the scene of the alleged shooting”United States v. Tibbs, No. 2016 CF1 19431, 2019 D.C. Super. LEXIS 9 (D.C. Super. Ct. Sept. 5, 2019); Missouri v. Goodwin-Bey, No. 1531-CR00555-01 (Mo. Cir. Ct. Dec. 16, 2016)
“qualitative opinions” can only be offered on the significance of “class characteristics”People v. Ross, 129 N.Y.S.3d 629 (N.Y. Sup. Ct. 2020)

The approach toward firearms testimony has evolved over the past two decades. The consensus approach, early on, was shared by a series of courts that adopted the formulation, “a reasonable degree of ballistic certainty.”324Diaz, 2007 U.S. Dist. LEXIS 13152, at *36; see also Commonwealth v. Pytou Heang, 942 N.E.2d 927, 945 (Mass. 2011); United States v. Simmons, No. 2:16cr130, 2018 U.S. Dist. LEXIS 18606, at *24–27 (E.D. Va. 2018); Cazares, 788 F.3d at 988; Monteiro, 407 F. Supp. 2d at 372; Taylor, 663 F. Supp. 2d at 1180; Ashburn, 88 F. Supp. 3d at 249; United States v. Hunt, 464 F. Supp. 3d 1252, 1262 (W.D. Okla. 2020). Thus, the court in Diaz allowed the examiner to testify “that cartridge cases or bullets were fired from a particular firearm ‘to a reasonable degree of ballistic certainty,’ ” as did a series of other federal courts.325Diaz, 2007 U.S. Dist. LEXIS 13152, at *36. In Monteiro, the district court ruled that the examiner could testify that the “class characteristics were in complete agreement,” but aside from observing that consistency, to a “reasonable degree of ballistic certainty,” no further probabilistic statement could be offered.326Monteiro, 407 F. Supp. 2d at 372. The court reasoned, “Allowing the firearms examiner to testify to a reasonable degree of ballistic certainty permits the expert to offer her findings, but does not allow her to say more than is currently justified by the prevailing methodology.”327Id. at 372.

It is not clear what a reasonable degree of certainty consists of—as a result, the U.S. Department of Justice has barred examiners in federal cases from using that or similar terminology:328U.S. Dep’t of Just., Uniform Language for Testimony and Reports for the Firearms/Toolmark Discipline Pattern Analysis 3 (2020).

An examiner shall not assert that two toolmarks originated from the same source with absolute or 100% certainty, or use the expressions ‘reasonable degree of scientific certainty,’ ‘reasonable scientific certainty,’ or similar assertions of reasonable certainty in either reports or testimony unless required to do so by a judge or applicable law.329Id. at 3.

The Department also barred examiners from making assertions of a “zero error rate” or infallibility.330Id. Those requirements marked a real charge from prior practice.

Second, during this time, some judges, like the Department of Justice itself, began to focus on the probabilistic claims by experts and limited toolmark experts’ testimony about conclusions that claim infallibility or the lack of any error rate—courts rejected assertions of zero error rates.331See, e.g., United States v. Romero-Lobato, 379 F. Supp. 3d 1111, 1117 (D. Nev. 2019) (acknowledging that the “general consensus” of the courts “is that firearm examiners should not testify that their conclusions are infallible or not subject to any rate of error, nor should they arbitrarily give a statistical probability for the accuracy of their conclusions”); State v. Terrell, No. CR170179563, 2019 Conn. Super. LEXIS 827, at *3 (Conn. Super. Ct. Mar. 21, 2019) (same); United States v. Glynn, 578 F. Supp. 2d 567, 574 (S.D.N.Y. 2008) (limiting testimony in part because when experts “make assertions that their matches are certain beyond all doubt, that the error rate of their methodology is ‘zero,’ ” there is a risk of “giving the jury the impression . . . that [the methodology] has greater reliability than its imperfect methodology permits”). Thus, courts rejected assertions of being “100% sure” or “certain.”332United States v. Parker, 871 F.3d 590, 600 (8th Cir. 2017). In Monteiro, the judge rejected the use of the phrase “a match to an exact statistical certainty.”333United States v. Monteiro, 407 F. Supp. 2d 351, 355 (D. Mass. 2006). Similarly, in United States v. Gardner, the judge held that the opinion could not be made with “unqualified” certainty.334Gardner v. United States, 140 A.3d 1172, 1184 (D.C. 2016).

A growing group of judges then offered intermediate approaches. Another District of Columbia judge held that an expert can testify that ammunition is “consistent with” being fired from the same firearm.335United States v. Sutton, No. 2018 CF1 009709, at *5 (D.C. Super. Ct. May 9, 2022) (permitting the examiner to opine “that the ammunition at issue is consistent with being fired from the same firearm”). The district court in United States v. Shipp ordered that the expert “may not testify, to any degree of certainty, that the recovered firearm is the source of the recovered bullet fragment or the recovered shell casing.”336United States v. Shipp, 422 F. Supp. 3d 762, 783 (E.D.N.Y. 2019). That court carefully examined the findings of the PCAST Report, and while it did not permit characterization of the level of certainty, the examiner could offer a statement of consistency.337Id. at 778. Other courts have taken this approach.338United States v. Davis, No. 4:18-cr-00011, 2019 U.S. Dist. LEXIS 155037, at *26–27 (W.D. Va. Sept. 11, 2019).

Going further to limit the testimony, in more recent cases, judges have barred any certainty-based statements at all. Thus, in the Tibbs ruling, the court held that the examiner could not offer any probability that the firearm in question could be included, but only that “the recovered firearm cannot be excluded as the source of the cartridge casing found on the scene of the alleged shooting.”339United States v. Tibbs, No. 2016 CF1 19431, 2019 D.C. Super. LEXIS 9, at *77 (D.C. Super. Ct. Sept. 5, 2019). In the Goodwin-Bey340State v. Goodwin-Bey, No. 1531-CR00555-01, slip op. at 7 (Mo. Cir. Ct. Dec. 16, 2016). ruling, the trial court did the same.341Id. (limiting testimony “to the point this gun could not be eliminated as the source of the bullet”). In United States v. Willock, the district judge ordered “a complete restriction on the characterization of certainty.”342United States v. Willock, 696 F. Supp. 2d 536, 546 (D. Md. 2010), aff’d sub nom. United States v. Mouzone, 687 F.3d 207 (4th Cir. 2012). Other courts have taken the same approach.343See United States v. White, No. 17 Cr. 611, 2018 U.S. Dist. LEXIS 163258, at *3 (precluding expert from testifying “to any specific degree of certainty as to his conclusion that there is a ballistics match”). Still other cases permitted the examiner to point to features and their similarities but not describe any level of agreement or consistency.344See, e.g., United States v. Green, 405 F. Supp. 2d 104, 124 (D. Mass. 2005); People v. Ross, 129 N.Y.S.3d 629, 642 (N.Y. Sup. Ct. 2020) (“The People may call an expert to testify as to whether there is evidence of class characteristics that would include or exclude the firearm at issue. . . . [T]he examiner may not opine on the significance of any marks other than class characteristics, as the reliability of that practice in the relevant scientific community as a whole has not been established. Moreover, any opinion based in unproven science and expressed in subjective terms such as ‘sufficient agreement’ or ‘consistent with’ may mislead the jury and will not be permitted.”).

None of these approaches adopt the approach of the American Statistical Association, which explains that to assert any degree of probability of an event, an established statistical basis must exist for that asserted degree of probability.345Am. Stat. Ass’n, Position on Statistical Statements for Forensic Evidence 2–3 (2019) [hereinafter ASA Report], https://www.amstat.org/asa/files/pdfs/POL-ForensicScience.pdf [https://perma.cc/T8EQ-BLZT]. Under this approach, an expert must be clear that no such statistical basis exists if none does exist.

The opinion that has gone farthest of all, however, is one of the most recent, a so-far unpublished opinion in 2023 by a trial judge in Cook County, Illinois. As noted, the judge wholly excluded firearms expert testimony, based on a review of scientific concerns with reliability.346See People v. Winfield, No. 15-CR-1406601, at 32–34 (Cir. Ct. Cook Cnty. Ill. Feb. 8, 2023).

2.  Limiting Non-Class-Based Opinions

Some jurisdictions under both Daubert and Frye have limited testimony to opinions offered on class characteristics only.347See, e.g., United States v. Adams, 444 F. Supp. 3d 1248, 1267 (D. Or. 2020); Ross, 129 N.Y.S.3d at 642 (“The People may proffer their NYPD ballistics detective as an expert in firearm and toolmark examination for the testimony on class characteristics as described above.”). That is, an expert can explain that the same type of gun fired the bullets or cartridge cases, but the expert cannot say that the same gun fired the bullets or cartridge cases. For example, here is the limiting instruction given by one federal judge who restricted the testimony to class characteristics:

[Firearm examiner’s] expert testimony is limited to the following observational evidence: (1) the Taurus pistol recovered in the crawlspace of [defendant’s] home is a 40 caliber, semi-automatic pistol with a hemispheric-tipped firing pin, barrel with six lands/grooves and right twist; (2) that the casings test fired from the Taurus showed 40 caliber, hemispheric firing pin impression; (3) the casings seized from outside the shooting scene were 40 caliber, with hemispheric firing pin impressions; and (4) the bullet recovered from gold Oldsmobile at the scene of the shooting were 40/l0mm caliber, with six lands/groves and a right twist.348Adams, 444 F. Supp. 3d. at 1267.

Courts have reasoned that descriptions of class characteristics are objective and measurable, whereas linking bullets to a particular gun is not “the product of a scientific inquiry,”349Id. at 1266. and “any opinion based in unproven science and expressed in subjective terms such as ‘sufficient agreement’ or ‘consistent with’ may mislead the jury and will not be permitted.”350Ross, 129 N.Y.S.3d at 642.

3.  Qualification and Proficiency Rulings

Judges have also focused on the proficiency of the particular expert to answer the preliminary question of whether a person is qualified to be an expert under Rule 702. Rule 702 requires that an expert witness have sufficient “knowledge, skill, experience, training, or education.”351Fed. R. Evid. 702 (requiring that an expert be “qualified as an expert by knowledge, skill, experience, training, or education”).

Typically, proficiency tests are administered by commercial test providers in which accredited labs are required to administer such tests annually.352Forensic Service Provider Accreditation, ANSI Nat’l Accreditation Bd., https://www.anab.org/forensic-accreditation [https://perma.cc/8GR7-4LPL]. For example, one leading provider, Collaborative Testing Services (“CTS”), makes available on its website the results of its tests for each discipline. CTS has cautioned that no “error rate” can be generalized from such tests because they are designed to be elementary.353Collaborative Testing Servs., Inc., CTS Statement on the Use of Proficiency Testing Data for Error Rate Determinations 3 (2010). Such tests are not proctored, can be taken in groups, have no time limit, include materials that are of unknown realism and difficulty, and are not “blind,” since participants know that it is a test.354Simon A. Cole, More Than Zero: Accounting for Error in Latent Fingerprint Identification, 95 J. Crim. L. & Criminology 985, 1029–30 (2005). However, the results do highlight the types of errors that practitioners may make. For example, a 2022 test included 7 participants, or 2% of the examiners, that failed to correctly identify the bullet that the known firearm had in fact fired in a test with a very small number of items; far higher numbers of examiners reported inconclusive responses which were also not accurate (but which CTS noted may follow lab practices).355See Collaborative Testing Servs., Inc., Firearms Examination Test No. 22-5261 Summary Report 3 (2022). CTS also noted that inconclusive responses were not counted as “outlier[]” errors, as “CTS is aware that many labs will not, as a matter of policy, report an elimination without access to the firearm or when class characteristics match.” Id.

In United States v. Cloud, the judge emphasized that one of the two examiners in the case had failed a proficiency test and was allowed to return to work after a second proficiency test, in which the examiner had to do an “in-depth consultation” with a supervisor.356United States v. Cloud, 576 F. Supp. 3d 827, 847 (E.D. Wash. 2021). The court found that it could not “in good conscience qualify [the examiner] as an expert with the requisite skill to perform fingerprint comparisons when her two most recent proficiency exams either contained an error or required a significant amount of assistance from her supervisor ” and further, the finding was bolstered by the portions of “testimony and performance reviews that touch on her skill, willingness to take correction, and confidence performing her work.”357Id. In the Willock case, the examiner’s “qualifications, proficiency and adherence to proper methods [we]re unknown.”358United States v. Willock, 696 F. Supp. 2d 536, 546 (D. Md. 2010).

Many courts traditionally focused on an expert’s credentials and self-professed expertise when conducting this inquiry into the qualifications of the witness.359See generally Brandon L. Garrett & Gregory Mitchell, The Proficiency of Experts, 166 U. Pa. L. Rev. 901 (2018) (arguing that objective evidence of proficiency, rather than credentials or self-professed expertise, should qualify experts). However, as one of the authors and Gregory Mitchell have argued, a careful inquiry into objective proficiency of the witness should be an integral part of the question whether a person should be qualified as an expert.360See id. at 940–49. Other courts have cited to the existence of proficiency testing as evidence of reliability, which as Garrett & Mitchell discuss, is not well supported. See, e.g., United States v. Johnson, No. (S5) 16 CR. 281 (PGG), 2019 U.S. Dist. LEXIS 39590, at *46 (S.D.N.Y. Mar. 11, 2019), aff’d, 861 F. App’x 483 (2d Cir. 2021) (“While these proficiency tests do not validate the underlying assumption of uniqueness upon which the AFTE theory rests, they do provide a mechanism by which to test examiners’ ability—employing the AFTE method—to accurately determine whether bullets and cartridge casings have been fired from a particular weapon.”). Indeed, such proficiency issues can raise larger red flags concerning the reliability of a crime lab unit and not just an individual examiner. Years before the Metropolitan Crime Lab had its accreditation revoked, as described in our introduction, a firearms examiner had failed a proficiency test after two colleagues had verified the work, implicating their own proficiency as well.361See Brandon L. Garrett, Autopsy of a Crime Lab: Exposing the Flaws in Forensics 94–95 (2021). Perhaps more careful attention to those proficiency tests could have prevented subsequent errors and systems failures of the firearms unit and the entire laboratory.

4.  As-Applied Challenges

Still additional challenges have focused on Rule 702(d), which used to provide that qualified expert testimony is admissible only when “the expert has reliably applied the principles and methods to the facts of the case.”362Fed. R. Evid. 702(d). These “as applied” challenges focus on the work that an expert does and not just whether they followed the right steps, but also whether their casework was actually supported by a valid method.363For a helpful explanation of what an as-applied challenge entails, see Edward J. Imwinkelried, The Admissibility of Scientific Evidence: Exploring the Significance of the Distinction Between Foundational Validity and Validity as Applied, 70 Syracuse L. Rev. 817, 832 (2020). Thus, some challenges have focused on, for example, the lack of documentation by firearms experts and the way they used their methods in a particular case.364For a case rejecting an as-applied challenge because the expert would not testify that a bullet came from a specific firearm, see United States v. Tucker, 18 CR 0119 (SJ), 2020 U.S. Dist. LEXIS 3055, at *3 (E.D.N.Y. Jan. 8, 2020). Some courts have found the presence of some documentation, such as “notes, worksheets, and photographs,” to be sufficient.365Ricks v. Pauch, No. 17-12784, 2020 U.S. Dist. LEXIS 50109, at *57 (E.D. Mich. Mar. 23, 2020); see also United States v. Harris, 502 F. Supp. 3d 28, 43 (D.D.C. 2020) (emphasizing that the expert shared “a description of his process and photo documentation.”); McNally v. State, 980 A.2d 364, 370 (Del. 2009) (finding cross-examination could adequately expose experts’ “lack of recollection” concerning application of methods).

III.  LESSONS FROM THE PATH OF FIREARMS EVIDENCE

The arc of judicial review of firearms evidence follows a pattern that is familiar in forensics more generally. Early judicial skepticism of a novel technique was overcome by claims of expertise relying on new technology (a microscope at the time), forceful claims to expertise by aggressive personalities (chiefly Major Goddard), some highly useful applications of the technique (to simply measure class characteristics), and steady accumulation of precedent. Then, as scientific critiques and evidence of error rates mounted, judges began to express some skepticism which has substantially increased in recent decades, producing a large body of law limiting firearms evidence in a range of ways.

That said, we underscore that other courts have not sought to introduce evidence concerning limitations of firearms evidence, much less imposed limitations. An appellate court in Missouri, for example, found no error in a judge’s refusal to allow defense attorneys to cross-examine the firearms expert concerning the findings of the NAS and PCAST reports.366State v. Mills, 623 S.W.3d 717, 729–31 (Mo. Ct. App. 2021), transfer denied (June 29, 2021) (“The trial court excluded the reports and their contents but did not deny defense counsel from asking questions about the flaws in toolmark and firearm examination as Appellant argues.”). Further, even in recent years, “many courts have continued to allow unfettered testimony from firearm examiners who have utilized the AFTE method.”367United States v. Romero-Lobato, 379 F. Supp. 3d 1111, 1117 (D. Nev. 2019) (citing David H. Kaye, Firearm-Mark Evidence: Looking Back and Looking Ahead, 68 Case W. Rsrv. L. Rev. 723, 734 (2018)).

The community of firearm examiners has mounted aggressive defenses of their work. In one memorable critique of how scientists and judges have raised questions concerning firearms comparison work, general counsel for the FBI wrote, “It is a lamentable day for science and the law when people in black robes attempt to substitute their opinions for those who wear white lab coats.”368Colonel (Ret.) Jim Agar, The Admissibility of Firearms and Toolmarks Expert Testimony in the Shadow of PCAST, 74 Baylor L. Rev. 93, 196 (2022) (“[C]ourts should recognize the long-standing reliability of the firearms identification discipline and the examiners who testify to that discipline.”). And yet it has been scientists—not judges—who have raised the deepest concerns about firearm examination. Statisticians, for example, criticize firearms comparison methods as having been “developed by insular communities of nonscientist practitioners” who, as a result, “did not incorporate effective statistical methods.”369William A. Tobin, H. David Sheets & Clifford Spiegelman, Absence of Statistical and Scientific Ethos: The Common Denominator in Deficient Forensic Practices, 4 Stats. & Pub. Pol’y 1, 1 (2017). As one litigator colorfully wrote in a Daubert brief, “Astrologers believe in the legitimacy of astrology. . . . And toolmark analysts believe in the reliability of firearms identification; their livelihoods depend on it.”370United States v. Cloud, 576 F. Supp. 3d 827, 844 (E.D. Wash. 2021).

The response to these scientific critiques has been to call them “flawed”371See Agar, supra note 368, at 166 (“Accreditation, widespread proficiency testing, the success of ATF’s NIBIN database, the Commerce Department’s recognition of firearms identification, and the reliance of the U.S. government on firearms identification to investigate and solve the assassination of a U.S. president serve as cornerstones for the ‘general acceptance’ of the firearms identification discipline.”). and double down on the claim that error rates are extraordinarily low. The FBI, for example, asserted in a 2022 case that there is an error rate of “1%.”372FBI, FBI Laboratory Response to Declaration Regarding Firearms and Toolmark Error Rates Filed in Illinois v. Winfield, May 3, 2022, at 3 (on file with authors). Federal prosecutors have repeatedly argued that “[f]irearms and toolmark identification meets all the Daubert criteria. Accordingly, there is no scientific or legal basis to exclude this evidence or even limit it.”373Gov’t’s Response to Defendant’s Motion in Limine to Exclude Ballistics Evidence, or Alternatively, for a Daubert Hearing at 23, United States v. Hunt, No. 5:19-cr-00073-R, 2020 WL 3549386 (W.D. Okla. April 27, 2020). Indeed, then-Attorney General Loretta Lynch more broadly responded to the PCAST report, upon its release, as not affecting the work of the Department of Justice: “We remain confident that, when used properly, forensic science evidence helps juries identify the guilty and clear the innocent. . . . While we appreciate their contribution to the field of scientific inquiry, the department will not be adopting the recommendations related to the admissibility of forensic science evidence.”374Gary Fields, White House Advisory Council Report Is Critical of Forensics Used in Criminal Trials, Wall St. J. (Sept. 20, 2016, 4:25 PM), https://www.wsj.com/articles/white-house-advisory-council-releases-report-critical-of-forensics-used-in-criminal-trials-1474394743 [https://perma.cc/XA3L-XHXE].

Some reactions in the field have been less defensive. Apparently in response to criticism by Judge Edelman, AFTE has opened its publications to outside viewing—one judge “applauds the publication’s changes and encourages AFTE and similar organizations to continue to open their publications up for criticism and review from the larger scientific community if they wish to meet Daubert’s rigorous standard.”375Cloud, 576 F. Supp. 3d at 842. However, the judge nevertheless found that the quality of the studies did not provide strong support for admissibility under Daubert.376Id.

One response by judges has been, as described, to limit the verbal formulations that firearms experts use when reaching conclusions. There are reasons to doubt that this compromise solution has been effective in communicating to jurors the limitations of firearms evidence. Two of us collaborated on a mock jury study examining how laypersons evaluate different firearms expert conclusions.377See Garrett et al., supra note 8. None of the limitations on firearms testimony adopted by courts, such as reasonable scientific certainty or more likely than not, had any impact on conviction rates except for the most far-reaching language, imposed in Tibbs, that barred any conclusion linking the firearms in question but rather permitting only a statement that a firearm cannot be excluded.378Id.

To be sure, the more recent rulings that permit only testimony concerning class characteristics go further than ruling out any language of inclusion. They limit the expert to testimony concerning objective measurements (for example, the width of the cartridge or bullet) and prevent more speculative testimony concerning probabilities that something came from a particular firearm. These rulings return firearms comparison to its roots: measuring objects. This can be useful and provide valuable information.

We have not seen judges take the approach to reliability, which is codified in Rule 702, that PCAST did, for example, insisting that “[t]he only way to establish the scientific validity and degree of reliability of a subjective forensic feature-comparison method—that is, one involving significant human judgment—is to test it empirically by seeing how often examiners actually get the right answer.”379An Addendum to the PCAST Report on Forensic Science in Criminal Courts 1 (Jan. 6, 2017).

In fact, some judges have expressly rejected this approach, stating that PCAST’s requirement of empirical study “goes beyond what is required by Rule 702.”380United States v. Harris, 502 F. Supp. 3d 28, 38 (D.D.C. 2020); see also United States v. Hunt, 464 F. Supp. 3d 1252, 1258 (W.D. Okla. 2020) (“[T]he Court declines Defendant’s invitation to restrict judicial review to techniques tested through black-box studies.”). However, there are strong reasons to think that jurors will benefit from more information regarding error rates and the reliability of the firearms comparison method, just as PCAST recommends and as mock jury experts have found productive—even just the bare acknowledgement that errors occur can impact jurors who assume that these experts are infallible unless told otherwise.381Brandon Garrett & Gregory Mitchell, How Jurors Evaluate Fingerprint Evidence: The Relative Importance of Match Language, Method Information, and Error Acknowledgment, 10 J. Empirical Legal Stud. 484, 503 (2013).

We note that this guidance extends not just to black box-type studies of the method, but also proficiency testing and other assessments of how well experts do their work in case-work settings, as well as blind testing, in which they do not know that they are being tested. Given how cognitive biases can impact the work of examiners in forensic settings, the evidence from black box studies may substantially underestimate error rates in actual casework.382See generally, e.g., Glinda S. Cooper & Vanessa Meterko, Cognitive Bias Research in Forensic Science: A Systematic Review, 297 Forensic Sci. Int’l 35 (2019). Moreover, jurors are extremely receptive to such information as well.383See generally, e.g., Gregory Mitchell & Brandon L. Garrett, The Impact of Proficiency Testing Information and Error Aversions on the Weight Given to Fingerprint Evidence, 37 Behav. Sci. & L. 195 (2019).

Nor have we seen judges take the approach of the American Statistical Association, which would require examiners to affirmatively state that there is no statistical basis for any probabilistic conclusion in their field.384ASA Report, supra note 345, at 4–5. Judges have, perhaps understandably, been far more comfortable with limiting conclusion language of experts than affirmatively requiring experts to explain limitations of their methods.

The 2023 amendments to Federal Rule of Evidence 702 encourage judges to more carefully consider that the proponent of an expert bears the burden to show that the various reliability requirements are met as well as that the opinions that the expert formed are reliably supported by the application of the methods to the data.385Committee on Rules of Practice and Procedure, June 7, 2022 Meeting 891–93, https://www.uscourts.gov/sites/default/files/2022-06_standing_committee_agenda_book_final.pdf [https://perma.cc/B8RW-YKCN]. That rule change, while reflecting prior law and not intended to change the substance of Rule 702, highlights the importance of judicial gatekeeping regarding the evidence that the proponent of the expert has that the work done, as well as the opinions reached, were grounded in reliable interpretation of data. The amendment supports the approach that we recommend: simply put, the exclusion of methods that are not demonstrated to be reliable. At a minimum, experts should also, as the American Statistical Association states, disclose all of the known limitations of their work.

Despite mounting scientific concerns and a limited response to the problem of firearms testimony by the Department of Justice,386We also note proposed standards from a different group that are in progress and largely restate the AFTE identification-based approach. See Firearms & Toolmarks Subcommittee, Standards: At an SDO for Further Development & Publication, Nat’l Inst. of Standards & Tech. (Mar. 1, 2022), https://www.nist.gov/osac/firearms-toolmarks-subcommittee [https://perma.cc/6ZWD-R7ED]. there has been a substantial federal investment in increasing the use of firearms comparison work. The federal database, the National Integrated Ballistic Information Network (“NIBIN”), has been supported by extensive federal grants, including regarding the expensive imaging equipment used on firearms evidence, to enter it into the database. Interestingly, the algorithms used to search that database remain a black box—the federal government has sponsored research on increasing the speed and efficiency of searches but not on how reliable “hits” are using the database.387Garrett, supra note 361, at 188. See generally William King, William Wells, Charles Katz, Edward Maguire & James Frank, Opening the Black Box of NIBIN: A Descriptive Process and Outcome Evaluation of the Use of NIBIN and Its Effects on Criminal Investigations (Oct. 2013), https://www.ojp.gov/pdffiles1/nij/grants/243977.pdf [https://perma.cc/2RJQ-5MGT].

Technology may eventually supply reliable means to provide quantitative information about the probability that a bullet or shell casing came from a particular firearm. Statistical approaches to this problem are under development, and one has been piloted by researchers with some promising initial results.388See CSAFE Develops New Bullet Matching Technology (Aug. 29, 2017), https://forensicstats.org/news-posts/csafe-develops-new-bullet-matching-technology [https://perma.cc/TD4H-QBZC]; Alicia Carriquiry, Heiki Hofmann, Xiao Hui Tai & Susan VanderPlas, Machine Learning in Forensic Applications, 16 Significance 29, 30–35 (2019). It may be that this is a scientific challenge that can be met. But for many decades, courts were willing to allow examiners to claim expertise that they lacked, based on assertions of experience, training, and proficiency that were not tested. Fortunately, now that those assertions have been minimally tested, some courts are stepping back to assess whether this expertise should be permitted. It is an object lesson in the acceptance and use of expert evidence in criminal courts, however, that has taken over a century for that shift to occur.

We end by emphasizing two other points. In this Article, we have focused on firearms comparison work, but it is only one specialty in the area of forensic toolmark comparison. It is among the most commonly used and has attracted sustained scientific and judicial attention, but as David Kaye and colleagues have importantly pointed out, “there is less research into the accuracy of associating impressions from tools such as screwdrivers, crowbars, knives, and even fingernails.”389Yale Law School Forensic Science Standards Practicum, Toolmark-Comparison Testimony: A Report to the Texas Forensic Science Commission 10 (2022) (“There are fewer limiting opinions involving source attribution to other tools, probably because fewer of these examinations are performed, and fewer reports bubble up to the courts.”). There is every reason to think that those other types of toolmark comparison raise similar or far larger reliability concerns.

Further, in this Article we have focused on criminal cases that proceed to a trial and evidentiary rulings at trial and on appeal. Yet, courts often do not have a Daubert hearing or issue written rulings regarding expert evidence questions.390United States v. Lee, 19-cr-641, 2022 U.S. Dist. LEXIS 150054, at *7 (N.D. Ill. Aug. 22, 2022) (“[S]ince the issuance of the NRC and PCAST reports, courts unanimously continue to allow firearms identification testimony.”). There have been high-profile wrongful convictions in cases involving firearms evidence, like that of Curtis Flowers who had six criminal trials, and no reported decisions discussing the firearms evidence involved.391See generally Jiaxin Zhu, Liangcheng Yi, Wenqian Ma, Ziyue Zhu & Guillem Esquius, The Reliability of Forensic Evidence: The Case of Curtis Flowers, Cornell U.L. Sch. Soc. Sci. & L., https://courses2.cit.cornell.edu/sociallaw/FlowersCase/forensicevidence.html [https://web.archive.org/web/20231014224452/https://courses2.cit.cornell.edu/sociallaw/FlowersCase/forensicevidence.html]. In a very interesting 2020 case, a judge found it appropriate for an exonerated person to introduce experts to show that the firearms evidence should have been exculpatory at the time of trial.392See generally Ricks v. Pauch, No. 17-12784, 2020 U.S. Dist. LEXIS 50109 *50 (E.D. Mich. Mar. 23, 2020) (denying defendant’s motion to strike plaintiff’s firearms experts). And most criminal cases are not tried. Lawyers may plea bargain cases based in part on the perceived power of a firearms comparison. Indeed, courts have regularly rejected application of Daubert reliability standards in other pretrial contexts, such as an application for probable cause relying on a firearms comparison.393See United States v. Rhodes, No. 3:19-CR-00333-IM, 2022 U.S. Dist. LEXIS 77231, at *16 (D. Or. Apr. 28, 2022) (“[P]robable cause in the context of a warrant is not subject to the Daubert standard.”). Further, laboratory audits have occurred based on revelations regarding errors in firearms work, which have not generated any written opinions in court, but which highlight the importance of forensic science commissions and other bodies tasked with investigating quality control failures in crime laboratories.394See generally, e.g., Tex. Forensic Sci. Comm’n, Final Report for Complaint Filed By Attorney Frank Blazek Regarding Firearm/Toolmark Analysis Performed At the Southwestern Institute of Forensic Science (April 2016), https://www.txcourts.gov/media/1440859/14-08-final-report-blazek-complaint-for-joshua-ragston-swifs-firearm-toolmark-analysis-20160419.pdf [https://perma.cc/EV5X-PC3M]; Justin Fenton, ‘Serious Questions’ Raised By Reports On Problems Inside Baltimore Police Crime Lab, Councilman Says, Baltimore Sun (Aug. 16, 2021, 2:18 PM), https://www.baltimoresun.com/news/crime/bs-md-ci-cr-crime-lab-folo-20210816-u6sbc72o25gjvfqeex4mfp2kvi-story.html [https://perma.cc/2VT4-8XX7]; Michigan State Police Forensic Science Division, Audit of the Detroit Police Department Forensic Services Laboratory Firearms Unit (2008). Thus, while there may be increasingly careful judicial review of firearms expertise in trial settings, much of the use of forensic evidence may remain largely unreviewed by judges.

CONCLUSION

We do not know how often people have been wrongly convicted based on erroneous firearms comparison conclusions. But we do know of people convicted based on firearms evidence testimony who have since been exonerated. For example, on January 16, 2019, Patrick Pursley was exonerated, in part because “evidence in 1993 was scant by today’s standards, and when you start with scant evidence you’re not in a good position to reevaluate it years later.”395Patrick Pursley, Other Murder Exonerations with False or Misleading Forensic Evidence, Nat’l Registry of Exonerations (last updated Feb. 27, 2022), https://www.law.umich.edu/special/exoneration/Pages/casedetail.aspx?caseid=5487 [https://perma.cc/E932-ACZR]. In that case, the judge found that defense experts demonstrated conclusively that the cartridge cases in question were not fired by the gun attributed to Pursley.396Id.

We have described how over the past one-hundred-plus years, judges’ initial skepticism of early firearms experts transformed into growing judicial acceptance, in large part because confident experts displayed new terminology, techniques, and technology like the comparison microscope. The result was—and still remains—“an overwhelming acceptance in the United States and worldwide of firearm identification methodology.”397United States v. Chavez, No. 15-CR-00285-LHK-1, 2021 U.S. Dist. LEXIS 237830, at *16–17 (N.D. Cal. Dec. 13, 2021). But despite a mountain of long-standing precedent, judicial acceptance of this testimony has eroded in recent years. After many decades of rote acceptance of the assumptions underlying the methodology, judicial interest in firearms expert evidence has exploded. Over half of the judicial rulings that we identified have occurred since 2009, the year that the NAS issued its pathbreaking report. Dozens of opinions limit testimony of firearms experts in increasingly stringent ways.

This sea change has occurred because of the work of lawyers, judges, and particularly scientists, who have played a key role in generating a new body of precedent. Scientists have demanded studies to examine questions of reliability, and they have exposed how the resulting studies uncovered deep concerns regarding error rates in firearms analysis. Firearms experts may have testified with confidence in the past. But today, they increasingly face defense experts who turn the microscope to the scientific flaws underlying firearms identification. In turn, judges have increasingly engaged closely with scientific research, error rate studies, and defense expert witnesses.

The Daubert revolution did not result in an immediate shift in how judges reviewed firearms evidence, but over time, judges have begun to grapple with the reliability standards. The scientific community continues to inform that work with detailed critiques. In turn, defense lawyers have launched more precise challenges that have shaped precedent.

The December 2023 revisions to Rule 702, designed to address both the burden to show that an expert is reliable and the manner in which experts reach and express conclusions, will solidify the focus—sharpened in firearms evidence rulings—on both of those important aspects of the judicial gatekeeping role. The resulting body of law has already reshaped how firearms evidence is received in criminal cases, and it provides important lessons regarding the slow, but perhaps steady, reception of science in our precedent-bound halls of justice.

APPENDIX

Appendix A.  Judicial Rulings Limiting Firearms Evidence, 2005–2022
CitationLimitation on Testimony
United States v. Felix, No. CR 2020-0002, 2022 U.S. Dist. LEXIS 213513 (D.V.I. Nov. 28, 2022)Limiting testimony to conclusions regarding class characteristics and whether individual toolmarkings were “consistent”
United States v. Stevenson, No. CR-21-275-RAW, 2022 U.S. Dist. LEXIS 170457 (E.D. Okla. Sept. 21, 2022)Limiting expert to “reasonable degree of ballistic certainty”
Winfield v. Riley, No. 09-1877, 2021 U.S. Dist. LEXIS 85908 (E.D. La. 2021)Limiting expert to “more likely than not” conclusion
United States v. Adams, 444 F. Supp. 3d 1248 (Or. 2020)Observational evidence permitted but no methods of conclusions relating to whether casings “matched” to be admitted
People v. Ross, 129 N.Y.S.3d 629 (Sup. Ct. 2020)Ruling that “qualitative opinions” can only be offered on the significance of “class characteristics”
United States v. Hunt, 464 F.Supp.3d 1252 (W.D. Okla. 2020)Permitting “reasonable degree of ballistic certainty”
State v. Raynor, 254 A.3d 874 (Conn. 2020)Permitting “more likely than not” testimony
United States v. Harris, 502 F. Supp. 3d 28 (D.D.C. 2020)Instructed expert to abide by DOJ limitations, including not using terms like “match” and not claiming to exclude all firearms in the world
Williams v. United States, 210 A.3d 734 (D.C. 2019)Finding error to permit expert to testify that there was not “any doubt” in conclusion
State v. Gibbs, 2019 Del. Super. LEXIS 639 (Del. Sup. Ct. 2019)May not testify to a “match” with any degree of certainty, and may not testify to a “reasonable degree” or “practical impossibility”
United States v. Tibbs, 2019 D.C. Super LEXIS 9 (D.C. Super. 2019)Limiting testimony to “the recovered firearm [that] cannot be excluded as the source of the cartridge casing found on the scene of the alleged shooting”
United States v. Davis, 2019 U.S. Dist. LEXIS 155037 (W.D. Va. 2019)Preventing testimony to any form of “a match”
United States v. Shipp, 422 F.Supp.3d 762 (E.D.N.Y. 2019)Preventing testimony “to any degree of certainty”
United States v. Medley, No. PWG 17-242 (D. Md. April 24, 2018)Permitting “consistent with” but no opinion fired by same gun
State v. Terrell, 2019 Conn. Super. LEXIS 827 (Conn. 2019)Prohibiting testimony regarding likelihood so remote as to be practical impossibility
United States v. Simmons, 2018 U.S. Dist. LEXIS 18606 (E.D. Va. 2018)Limiting to “a reasonable degree of ballistic . . . certainty”
United States v. White, 2018 U.S. Dist. LEXIS 163258 (S.D.N.Y. 2018)Holding that expert may not provide any degree of certainty unless pressed on cross-examination and may then present “personal belief”
State v. Jaquwan Burton, Superior Court, No. CR14-0150831 (Conn. Super. Ct. Feb. 1, 2017)Permitting “consistent with” but no opinion that it was fired by same gun
Missouri v. Goodwin-Bey, No. 1531-CR00555-01 (Mo. Cir. Ct. Dec. 16, 2016)Limiting to “the recovered firearm [that] cannot be excluded as the source of the cartridge casing found on the scene of the alleged shooting”
Gardner v. United States, 140 A.3d 1172 (D.C. 2016)Error to admit “unqualified” testimony with “100% certainty”
United States v. Cazares, 788 F.3d 956 (9th Cir. 2015)Limiting to “reasonable degree of scientific certainty”
United States v. Black, 2015 U.S. Dist. LEXIS 195072 (D. Minn. 2015)Limiting to “reasonable degree of ballistics certainty” and barring “certain” or “100%” conclusions
United States v. Ashburn, 88 F. Supp. 3d 239 (E.D.N.Y. 2015)Limiting to “reasonable degree of ballistics certainty” and precluding “certain” and “100%” sure statements
United States v. McCluskey, 2013 U.S. Dist. LEXIS 103723 (D.N.M. 2013)Limiting testimony to “practical certainty” or “practical impossibility”
United States v. Mouzone, 687 F.3d 207 (4th Cir. 2012)Approving trial ruling limiting any expression of certainty
United States v. Love, No. 2:09-cr-20317-JPM (W.D. Tenn. Feb. 8, 2011)Barring testimony of “practical” or “absolute” certainty
Commonwealth v. Pytou Heang, 942 N.E.2d 927 (Mass. 2011)Limiting to “reasonable degree of ballistics certainty”
United States v. Cerna, 2010 U.S. Dist. LEXIS 144424 (N.D. Cal. 2010)Limiting to “reasonable degree of ballistics certainty”
United States v. Willock, 696 F. Supp. 2d 536 (D. Md. 2010)“[A] complete restriction on the characterization of certainty” and precluding “practical impossibility” conclusion
United States v. Taylor, 663 F. Supp. 2d 1170 (D.N.M. 2009)Limiting to “reasonable degree of scientific certainty”
United States. v. Glynn, 578 F. Supp. 2d 567 (S.D.N.Y. 2008)Limiting to “more likely than not”
United States v. Diaz, 2007 U.S. Dist. LEXIS 13152 (N.D. Cal. 2007)Limiting to “reasonable degree of certainty in the ballistics field” and no testimony “to the exclusion of all other firearms in the world.”
United States v. Monteiro, 407 F. Supp. 2d 351 (D. Mass. 2006)Limiting to “reasonable degree of ballistic certainty”
Commonwealth v. Meeks, 2006 Mass. Super. LEXIS 474 (Mass. Super. Ct. 2006)Requiring examiner to present “detailed reasons” for rulings
United States v. Green, 405 F. Supp. 2d 104 (D. Mass. 2005)Barring “to the exclusion of all other guns” language
97 S. Cal. L. Rev. 101

Download

* Neil Williams, Jr. Professor of Law, Duke University School of Law, Faculty Director, Wilson Center for Science and Justice. Many thanks to Anthony Braga, Mugambi Jouet, Daniel Klerman, Charles Loeffler, Thomas D. Lyon, Aurelie Ouss, Danibeth Richey, Greg Ridgeway, D. Daniel Sokol, and the participants at workshops at University of Southern California Gould School of Law, a Center for Statistics and Applications in Forensic Evidence webinar, and the Department of Criminology, University of Pennsylvania for their feedback on earlier drafts, to Stacy Renfro for feedback on the firearms case law database, to Richard Gutierrez for helpful comments, and to Hannah Bloom, Erodita Herrera, Megan Mallonee, Linda Wang, and Grace Yau for their research assistance. This work was funded (or partially funded) by the Center for Statistics and Applications in Forensic Evidence (CSAFE) through Cooperative Agreements 70NANB15H176 and 70NANB20H019 between NIST and Iowa State University, which includes activities carried out at Duke University and University of California, Irvine.

† J.D., Duke University School of Law.

‡ Visiting Research Professor of Law, University of Southern California Gould School of Law. Professor of Psychology and Criminology, University of California, Irvine.