Scott Alexander posted an insightful (and blessedly short—thank you!) article about how most ideas for solving debate are doomed to fail. He nevertheless concedes that it is a worthwhile goal. (I agree!)
He is probably right! It is such an absurdly large goal that it seems impossible, possibly because it is impossible. But I believe that there is plenty of opportunity to systematically improve discourse for society at large. And although most people agree this is a worthwhile goal, I think the stakes may be higher than that.
If you haven’t read his article, or subscribed to Astral Codex Ten, you should do that first.
Scott’s skeptical argument
I will go over each of the main thrusts of Scott’s arguments here, show a representative quote, and add my interpretation. Although I agree with many or all of Scott’s points, I nevertheless come to a different conclusion. Perhaps that itself is the best evidence that Scott is correct.
“Debate” almost never corresponds to mappable arguments.
The simplest “solve debate” proposal is the argument map. Some technology helps people decompose arguments into premises and conclusions, then lets skeptics point out where the premises are wrong, or where the conclusion doesn’t follow from the premise.
But almost no real argument works that way. Even in the best-case scenario, where an argument almost works that way, it doesn’t really work that way. Suppose you’re having an argument about COVID lockdowns. Someone says “lockdowns hurt the economy”. Now you’re stuck in a giant fight about whether that claim is true (…) .
Scott’s point here is that argument maps make things more complicated, not easier, and I am pretty sure I am in agreement. He goes on to elaborate that most arguments are too complex to be compatible with syllogism, so you cannot reliably make inferences based on logic without flattening the argument beyond usefulness.
Eventually he concludes that following an argument map wouldn’t so much help people understand as encourage one side or the other to point at one particular premise, declare it either right or wrong (based on the argument map), and claim victory for their side. Nobody gets smarter, and aggregate knowledge does not expand.
After briefly testing out argument maps on a subject I’m familiar with, I have to say I agree with Scott that argument maps do not seem to be that useful. In practice, they seem to rush you to a conclusion that you don’t understand yourself. But they do, at least, have the benefit of documenting sides of the debate, which I think could be useful for somebody investigating a case they don’t already know.
Arguments rarely hinge on one person being simply wrong and stupid
But arguments rarely hinge on false facts. Even an insane person like Alex Jones rarely says specific false facts. In the few cases where there are specific false facts, most people, when forced, are happy to jettison the false fact and continue making the argument on other grounds.
And arguments rarely hinge on specific named fallacies. Even when they do, those fallacies often provide some kind of useful information (it’s technically ad hominem fallacy to say that Alex Jones is wrong about Sandy Hook because he’s a lying loon, but I actually know very little about his Sandy-Hook-related arguments, and I think relying on his general mendacity and looniness is a useful proxy here).
For sure, Scott is onto something here.
To use another example, suppose you want to prove to loved ones that they belong to a cult so that you can get your family back. You may have an insane amount of evidence that they follow false prophets, and worship an idol they believe is Jesus Christ, but pointing these things out counterintuitively seems to harden their hearts because they feel attacked.
This is a high-stakes situation where the truth would set them free, unite families that are divided by lies, save countless children from child abuse, and open up free speech among one of the most repressed cohorts that exists.
Sadly, in most cases, we cannot save people so utterly deceived. Most will cling to their beliefs in spite of contrary evidence. But we certainly can do our best to expose the truth for those whose lives are not already committed to lies.
The hardest problem for any social technology is getting users
We’ve talked about this before for dating apps. And if it’s that hard to lure people in with the promise of sex, it’s probably even harder to lure them in with the promise of logical accuracy.
Like dating, arguments require two compatible people. That means your app is useless until it has a big enough population that, whenever someone wants to argue, there will be a compatible person willing to take the other side of that argument in a reasonable amount of time. That’s hard to bootstrap.
There is no doubt that this is true. A long time ago Google created their version of Facebook, something called Google+, and it failed miserably despite integration into products with millions in its user base. The predecessor, Orkut, arguably enjoyed more relative success (in certain markets) than Google+ ever achieved. The products were not bad, they just weren’t needed. Why post on Google+ when you’re already on Facebook, where everybody else was (at the time)?
Here’s Scott again:
You might think: “But don’t people like arguing on the Internet?” No. Look closer. People like taking drive-by potshots on the Internet - retweeting some link that makes them feel like they’ve successfully embarrassed their ideological enemies.
That quote above, might as well be a reference to me, because I am exactly that kind of a petty know-it-all. It’s why I’ve largely retreated from social media, even Substack: because I can’t trust myself to act in good faith.
And one final gem of a quote:
Even I don’t like arguing on the Internet. I like writing a post saying that my opinion is right. Sometimes people tell me that actually, my opinion is wrong. This frustrates me. I feel obligated to respond, but the response is an unfortunate step between me voicing my opinion and everybody agreeing that my opinion is right.
Yes, frustrating indeed. Here again, Scott names exactly my problem. I want to go out and tell the world that, no, actually I have the answer to this, that, and the other problem, but unfortunately that requires me to actually write something to persuade. It is a lot of work. Work is not a problem, but it is a constraint.
I’ll tell you this right now: Claude Code makes it easier to build an app than write an essay. (I am unwilling to let Claude write for me except for documentation, and where time constraints trump originality concerns.)
For example, Scott has another excellent piece on inflation versus perception that I’ve been meaning to respond to. For such a high-quality piece of writing, my response would necessarily have to refute the CPI as an acceptable measurement of inflation, and then propose a better one, or at least some standard a better one must meet.
So I must concede that there is no system that can compensate for humans whose only mode is sloppy potshots that, in the best case, vaguely gesture toward a point they fail to articulate themselves. That is my failure, and it is the one that I am trying to avoid in the future by building a reliable research tool that I can use to improve my understanding of a given topic. If it can be useful to other users, that’s a blessing I’m not counting on.
This hasn’t worked in two thousand years of arguing
Most dating apps are doomed. But one reason for optimism about dating apps is that people have made them before. They’ve been known to occasionally work. And they’re an extension of things like matchmaking resumes and classified ads which have worked for centuries. What’s the argumentative equivalent?
I will concede that nothing as audacious has ever been attempted before, and the world is right to be skeptical. And I will concede that there is every reason to believe this is indeed impossible, a fool’s errand. But I have good reason to disagree, and I will lay out my argument later.
Other skeptical arguments
I’m no stranger to skepticism about this idea. In the rare case that somebody actually understands my goal—and how could they when I fail to articulate it?—their curiosity is quickly surpassed by their incredulity. If they are my friend, then they offer kind guidance to give up now before wasting my time. If they are not my friend, the advice is more blunt, but similarly skeptical. When friend and foe offer me the same advice, I should certainly listen, shouldn’t I?
Most of the objections are pretty standard:
Nobody owns the Truth (with a capital “T”), and you are not the exception.
Won’t people just use the tool to lie?
Won’t people just continue to ignore facts?
Won’t people just stick to their own uninformed bubbles?
This is, for many good reasons, impossible.
Here is a clip of a fairly representative reaction to this idea from Matt Welch and Michael Moynihan of The Fifth Column, Special Dispatch #77, Luscious Scumbag Lips. I fully understand their skepticism, just as I do from Scott Alexander, and everybody else.
Matt Welch: This is Bryan. It’s kind of a weird one, but kind of interesting. Howdy, guys. Bryan with a Y. We’ve got a few of those, too.
Matt Welch: This message is super overdue. First off, you guys are awesome. Miami was awesome. And after today, I’ve confirmed that flying first class is awesome. Okay.
Michael Moynihan: Yeah, yeah, yeah.
Matt Welch: You’re undoubtedly salivating at the prospect of me upgrading my subscription, but I assure you, I will continue flying coach. Bryan! Now onto my question. Many, many Patreon episodes ago, you guys talked about journalism for real life and being smart consumers of information. This is an idea which fascinates me and which I think could be helped by decentralized software, similarly to Wikipedia.
Matt Welch: Wouldn’t it be cool if we could use software to mitigate the problems of partisan bias, incompetence, or even malice? That’s not nice to say about Michael. Since you guys have loads of experience as editors, I’m wondering if you currently, or have ever, used software to keep track of fact checking, context checking, and evaluation of the overall thesis of a given article? If so, I’d love to hear about it. If you’re wondering what I’m drinking, it’s just a standard Woodford and Coke, but it tastes better when it’s included in your flight.
Matt Welch: Thank you all for the stellar work you do. I appreciate you. Looking forward to the day you visit Seattle or the Phoenix area, where I’m moving next month. We will visit both places, Bryan. Moynihan, you know software. You’ve worked with software. Are you software?
Michael Moynihan: Well, I don’t know if there is any software that would actually work in that kind of context. Yeah, I see what he’s saying, but I think it’s probably an impossibility.
Thanks again to Matt Welch and Michael Moynihan for answering me. I am grateful! And again, I understand the skepticism. I would think the same thing if I were you.
Since this is from a paid episode of The Fifth Column, I am keeping the audio clip short, but you should definitely just go subscribe.
Here is Claude’s summary of the full 11-minute answer:
The question. A listener named Bryan, who mentions moving to the Phoenix area, asks whether decentralized, Wikipedia-like software could help mitigate partisan bias, incompetence, or malice in journalism — and whether the hosts, as working editors, have ever used tools to track fact-checking, context, and the soundness of an article’s overall thesis.
Moynihan’s answer: probably impossible. The hard part isn’t the tooling, it’s that “truth” is elastic and every faction claims to be defending it — school board fights over American history, arguments about bias in political coverage. He grants one measurable exception: story selection bias. You could quantify, say, how much column space Scandinavian papers give Israel relative to other conflicts. But as Welch points out, that’s a research finding, not a consumption tool for readers.
Welch’s answer: it starts at home. He recalls Michael Kinsley’s “wikitorial” experiment at the LA Times as a well-intentioned gimmick that didn’t work. The most software can do is grade seriousness, veracity, and partisan lean — which people already attempt. He’s skeptical of NewsGuard and Steve Brill’s version of this, citing Phil Magness’s criticism that NewsGuard had to walk back its own pronouncements as the lab leak theory and other COVID claims shifted. Welch also admits his own reading has degraded under social-media-driven consumption, and expects media trends to make that worse.
COVID as the test case. Both use the pandemic to argue the point: guidance changed daily (ventilators, rubber gloves), so real-time “disinformation” labeling was inherently unstable. They cite the denunciation of Stanford researchers — Ioannidis and Bhattacharya (garbled in the transcript) — who were accused of killing people for asking questions, and John Tierney’s City Journal piece on dissenting researchers losing publication access. Their distinction: good-faith researchers who get things wrong aren’t the same as Plandemic-style conspiracy content, and appointing guardians of a truth that’s still developing produces witch hunts rather than accuracy.
Clearly there is a long and terrible history of fact-checking and “truth” products failing. Some of them, like Snopes, were excellent for years before ultimately succumbing to partisan bias. Others, like NewsGuard, failed out of the gates for a variety of reasons. I don’t use those products, editorial or otherwise, because they are like fallible prophets: useless.
Other products, like Ground News, succeed at a much narrower scope. As Michael Moynihan noted, selection bias is a real use and quantifiable use case. I like Ground News just fine, but its narrowed scope makes it inherently less interesting.
I must concede these points up front:
If somebody is determined to avoid the truth, they will be successful.
If somebody is determined to lie, nobody can stop them.
If somebody claims omniscience, they are dishonest.
Nevertheless, there is opportunity for improvement.
Where did previous attempts fail?
History is a graveyard littered with failed attempts to solve for truth at scale, more than enough to justify skepticism. Even from huge companies with money to burn.
The most important thing is to acknowledge the reasons for prior failures. Namely, nobody knows everything about everything. X-Ray cannot identify Truth®, only help to identify alternate perspectives. But maybe that is enough.
Compared to the graveyard of past attempts, I think it seems most closely aligned with Truth Goggles. From the project website:
Truth Goggles attempts to decrease the polarizing effect of perceived media bias by forcing people to question all sources equally by invoking fact-checking services at the point of media consumption. Readers will approach even their most trusted sources with a more critical mentality by viewing content through various “lenses” of truth.
How is X-Ray different?
X-Ray is built on top of a NOSTR (Notes and Other Stuff Transmitted by Relays), a decentralized social media protocol with built-in Bitcoin’s Lightning Network built-in. Whenever you publish to NOSTR, you’re actually publishing a JSON object with a complete schema to describe the data. There is an accompanying work-in-progress NOSTR Implementation Possibility (NIP) draft to go along with X-Ray.
The most important thing that NOSTR enables is a standardized base data layer across all users and clients, including references to the same entities. Think of a giant database that anybody can read and write to at will, without asking for permission. This means that all of the work that captured by X-Ray (and future apps) can accrue, and be used by many people, including passively by those are who completely uninterested in doing their own research. That is one possibility that can be engineered by building on top of NOSTR.
How does it work?
X-Ray is a manifest-v3 browser extension that anybody can install from Github right now, and soon from the relevant extension stores. It is still rough around the edges, and not ready for wide release. There are many bugs to slay, and interfaces to clean-up. But it works, and you can use it now if you’re willing to put up with bugs.
You will need your own API keys for the following:
Anthropic - Provides most LLM features on the backend.
AssemblyAI / Deepgram - Provides transcription for audio and video files, with voice diarisation.
HuggingFace - For local LLM voice transcription through pyannote.
Here’s the basic workflow:
Open X-Ray and create a case, eg, “What is the origin of Covid?”
Visit and capture articles, videos, and podcasts relevant to the case question.
Use Suggest feature to have Claude identify persons, places, things, organizations from the article, and the claims associated with them.
This step also evaluates each claim against the specific case question, and caches the result for later case synthesis.
From “My Archive”, navigate to the “Cases” tab, and select the “Dashboard” button to open your case dashboard.
Scroll down and select “Analyze Corpus”.
This step synthesizes case-related claims from across the entire corpus to produce a Summary, Cruxes of Disagreement, Load Bearing Claims, and Coverage Gaps.
Download or publish the Case Brief output.
There are lots of other experimental features inside of X-Ray, but this is the only part that is well-tested and confirmed high-quality.
Primary use cases
Personal research tool - X-Ray demonstrates value a N=1 users by producing thoroughly researched case briefs. If nobody else ever uses it, it is already a worthwhile tool.
Passive research consumption - Downstream information consumers can benefit from work that a relative few power users produce, simply by visiting a URL with metadata associated. Consider YouTube videos that make claims that have already been debunked elsewhere. If having access to additional knowledge costs nothing, why wouldn’t a user take advantage of that?
Journalism reputation scoring - As data accrues over time, it will be possible to aggregate scores of specific journalists and organizations across various dimensions.
Infinite future use cases - What can you do with a giant decentralized worldwide database? Lots of things, even dating apps. And many others I’ve never imagined.
Example case briefs from X-Ray
The following are actual research cases produced by X-Ray, and published to NOSTR.
Final note on the importance of solving debate
The United States have ceased to be united in practice. Child abusers are protected by corrupt religious leaders. Free speech is under attack, and more expensive than ever. Political violence is tragically commonplace. Anti-semitism is once-again rearing its ugly head. The office of the President of the United States is nakedly corrupt. Men and women hate each other. Socialism is back from the dead, and totalitarianism cannot be far behind. How long until civil war breaks out?
None of these can be solved my humble tool, but the truth helps all who receive it, even if it is painful at first.
Communication is important, and effort is critical
For many reasons, I have chosen to remain silent for a long time.
Part of that was due to procrastination. Articulating arguments is a skill that I haven’t exercised since college, and one I never really respected. So that part of my brain has significantly atrophied, and I am not a fan of it.
Part of that was due to busyness with other things, like trying to understand why my family threw me away like so much trash after I complained about being abused as a child, and demanded respect for the first time in my life. Narcissism breeds narcissism, and I may be a recovering narcissist myself.
Part of my silence was the sheer number of subjects that I need to cover. There’s a lot, and none of them are simple. Prioritization has been genuinely difficult, and exacerbated by my desire to be a source of 100% signal and 0% noise. That is, again, impossible, but something I aspire to do. The advent of AI is a full-blown miracle, and finally allowed me to develop X-Ray after years of thinking about it.
Even as I write this, I am two-hours past the 11:59pm 8/31/26 deadline I set for myself. And the post is still rushed and sloppy. Oh well, publishing a good faith effort is better than never publishing at all.
Only forgiveness can resolve our differences
Last Sunday, I watched an excellent sermon by RT Kendall on “Total Forgiveness”. For the first time in years, I was able to let go of the burden of my anger toward a large roster of family, coworkers, religious leaders, and former friends.
Each of us requires mercy and forgiveness on a daily basis. To think otherwise is, quite frankly, arrogant. And I confess: I have been quite arrogant.
When I consider how much I rely upon the grace and forgiveness of others, I am humbled by that mercy. I recognize that I will need it now until the end of eternity, and therefore I must also forgive. It is not easy, but it is worth it.
And it is humanity’s only chance.












