THE real test of AI is not what it knows, but how it judges. Now what happens when an AI system is given better evidence and still reaches the same unfair conclusion? That question has become more urgent as AI systems increasingly stand between information and judgement. We tend to assume that if a model can search reliable documents before answering, its decisions will improve. But this confidence may be misplaced.
Most people understand the basic appeal of retrieval-based AI. Instead of asking a language model to rely only on what it learned during training, we first give it relevant material. A system judging a country’s climate record, for instance, can be shown official climate pledges and international agreements. The expectation seems reasonable: better evidence should lead to better judgement.
Yet judgement is not the same as fact-finding. A model can locate a date or a treaty clause with precision. But deciding whether a country has acted responsibly on climate change requires something more. The model must interpret evidence, weigh it against an external standard and revise whatever assumptions it carries. That last step is where the real difficulty appears.
Consider climate governance. Large language models can be asked to assess countries having very different economic circumstances and levels of vulnerability to climate change. Their judgements can then be compared with independent assessments such as those produced by the Climate Action Tracker, which evaluates national climate action using policy and emissions evidence. The models can also be supplied with official climate material, including international conference decisions and country-specific commitments.
Good evidence is essential, but not enough.
The surprising result is not that the models sometimes err. That is not out of the ordinary. The fact is that improving the evidence does not reliably improve judgement. Even when the information is relevant to the country being assessed, the final verdict may be no closer to the independent benchmark. More revealingly, the models react to what they are shown. Their answers often change after receiving evidence. But a changed answer is not necessarily a better one. Helpful revisions can be cancelled out by harmful revisions. The system is reading the material, yet it is not consistently converting that material into sounder judgement.
This distinction goes far beyond climate policy. Imagine an AI system assisting in university admissions, recruitment, credit decisions, or public policy. An institution may proudly say that the model is connected to verified records and current documents. That sounds reassuring. But access to reliable information does not guarantee that the model will interpret it fairly. A system can possess good evidence, while continuing to filter that evidence through assumptions formed during training.
Climate governance makes this problem concrete. Language models can still place undue blame for climate inaction on vulnerable countries, even after accounting for variations in their actual climate performance. As a result, nations that have made a comparatively small contribution to the historical issue could face harsher criticism than their independent climate data support.
This shouldn’t feel abstract to Pakistanis. Countries exposed to floods, heatwaves, and other climate shocks already have to argue their case in global systems shaped by unequal power. If automated systems begin to support decisions about climate responsibility, finance, or credibility, an old inequality can acquire the new appearance of neutrality. A machine-generated judgement may look objective because it is produced by software and supported by documents. That appearance can be dangerously persuasive.
A key takeaway is that missing information isn’t necessarily the main issue. A model may occasionally have a systemic tendency that fresh data finds difficult to overcome. Judgement can improve when that tendency is rectified and measured against external benchmarks rather than when better recovered material is provided. This is not a panacea. But it does let us know where the true vulnerability might lie.
Hence, the debate on responsible AI must go beyond asking ‘what information did the model receive?’ We should also inquire as to how such information influenced its assessment and whether the outcome was more in line with a reliable external norm. Before such systems are trusted with important judgements, this distinction should influence how they are tested. Good evidence is essential. It is simply not enough. As AI enters decisions about countries, institutions and people, we should resist the comforting belief that connecting a model to better documents automatically makes it fair. Information can reach the machine and still fail to reach judgement.
The writer is an AI engineer and researcher.
Published in Dawn, October 6th, 2026



























