Shann Memorial Lecture, University of Western Australia, Perth

Australian Treasury

1. How many hoops should a jobseeker have to jump through?

I acknowledge the Whadjuk Noongar people, on whose lands we meet, and recognise all First Nations people present. Thank you to the University of Western Australia, to Alison Preston, president of the Economic Society of Australia (Western Australia), and to everyone involved in organising the Shann Memorial Lecture. It is a huge honour to give a lecture bearing Ed Shann’s name. I will say more about Shann shortly, and how his approach to economics fits today’s theme.

But I want to begin with a much more recent Australian story about how our government is using randomised trials to shape government policy.

When people receive income support through Workforce Australia Online, the government sets mutual obligations. Jobseekers can meet these obligations through job applications, training, study and other activities. Each activity is allocated a certain number of points under the Points Based Activation System. The standard monthly target has been 100 points, which typically includes a minimum of 4 job applications.

The logic seems straightforward. Looking for work takes effort. Requiring more effort should, up to a point, increase the chances of finding a job.

But where is that point?

How many applications should someone be required to submit? How demanding should the points target be? At what stage does an obligation stop encouraging useful job search and begin consuming time that could be spent finding work in other ways?

Governments often answer questions like these through judgement. Officials study the evidence. Ministers weigh competing arguments. Eventually, a target is chosen.

In this case, the Australian Government ran an experiment (Department of Employment and Workplace Relations, forthcoming).

Almost 40,000 job‑ready jobseekers were randomly divided into 3 groups. One remained under the existing mutual obligation arrangements, with a target of 100 points each month. For the other 2, the requirements were reduced to 70 points per month (including 3 job applications) or 50 points per month (including 2 job applications).

Six months later, the employment results were strikingly similar. The 3 groups had worked an average of 10.4 weeks, 10.3 weeks and 10.5 weeks. Their earnings and income support payments were also very similar.

Randomised trials don’t just produce statistics – they can be used to garner qualitative data too. Interviewing participants from each group, the researchers also found that people facing lighter requirements tended to report less stress. Some said the change gave them more time to pursue activities they thought would actually help them find work.

Before the trial, there were plausible arguments in both directions. Tougher requirements might strengthen the incentive to get a job. Lighter requirements might allow people to search smarter, focusing on the particular job where they were most likely to get hired.

The experiment allowed government to find out. It provided evidence that helps us design a system that supports people move into work.

That is the subject of today’s talk.

We generally judge governments by the quality of the decisions they make. I want to suggest that we should also judge them by the systems they build for learning what happens afterwards. In an uncertain world, policy won’t always get it right the first time. The key is to build a feedback loop that ensures our policies become a little better every year. Policy by evidence, not gut feel.

There are 4 parts to this this lecture.

First, I will look at the economics of experimentation: the option value of learning, the trade‑off between exploration and exploitation, Bayesian updating, and the value of maintaining a portfolio of policy approaches.

Second, I will look at what Australia is already learning from randomised trials, from banking competition to electric‑vehicle charging, charities and parental leave.

Third, I will turn from individual experiments to the institutions that make learning routine: the Australian Centre for Evaluation, linked microdata, recent Commonwealth evaluation reforms, and emerging approaches in Britain and Canada.

Finally, I will ask what federalism and lotteries might contribute. In our federation, 8 states and territories generate policy variation every day. With better evaluation architecture, some of that variation can become a source of knowledge. Likewise, lotteries can serve a dual purpose: a fair way of allocating scarce resources, and an opportunity to discover what works.

Government productivity depends partly on how efficiently the state delivers what it has chosen to do. It also depends on how quickly the state learns whether it chose well.

2. The economics of finding out

Economists have several ways of thinking about the value of learning.

The first is the idea of option value: the benefit of keeping your options open so you can make a better decision later when you have more information. When a decision is hard to reverse, uncertainty is expensive. Committing early can lock in a weak policy for years. A trial preserves the ability to change course. We accept some delay in exchange for information that may improve policy in the long run.

This is familiar in investment. A firm deciding whether to build a factory may value the option to wait until demand becomes clearer. Governments face a similar calculation. If a program can be tested on 20,000 people before being extended to 2 million, the experiment itself has an economic return. Its value rises with the scale of the eventual program and the uncertainty around its effects.

A second idea is the exploration‑exploitation trade‑off. Imagine choosing repeatedly between restaurants. One strategy is to always stick with your favourite eatery. Another is to try that place your foodie friend won’t stop talking about. The first exploits current knowledge. The second explores.

Public policy faces the same problem. Once an approach looks promising, there are good reasons to extend it. Yet a system that always makes the same choice risks failing to improve.

Randomised trials allow us to rigorously evaluate alternatives (Leigh 2018). They are like a pilot program, but with a control group. The benefit to taxpayers from better policies are almost invariably larger than the cost of running the trial.

A third idea is Bayesian updating. We begin with beliefs about how the world works. New evidence changes those beliefs according to how informative the evidence is. Strong evidence should move us further than weak evidence.

That’s sound science, but things can be trickier in a big organisation. Programs accumulate budgets, contracts, IT systems and established delivery routines. Changing direction can be costly. A learning state therefore needs institutions that make updating easier, so that fresh evidence feeds back into subsequent decisions.

The fourth idea is portfolio theory. Investors diversify because uncertain returns need not move together. Governments can gain something similar from policy diversity. When different jurisdictions or service providers try different approaches, the system avoids staking everything on one theory of the world. This diversity can be valuable, provided that it is part of a learning process, where the variation creates information for policy makers. In particular, it is essential that outcomes are measured carefully and comparisons are credible.

In short, governments should value information about what works, and governments should design programs in such a way as to generate more of it.

3. Ed Shann: Intellectual restlessness as a public virtue

Ed Shann is an apt figure for a lecture about learning.

He arrived at the University of Western Australia in 1912 as its foundation professor of history and economics. He later served as vice‑chancellor, helped establish economics and economic history here, and taught students who would go on to shape Australian public life, among them H. C. Coombs and Arthur Tange. His students remembered a teacher who conveyed the excitement of intellectual exploration alongside a strong sense of social responsibility (Snooks 1988).

Like many of us, Shann’s views evolved over time. As a young scholar in Britain, he encountered the Fabians, including Sidney and Beatrice Webb and George Bernard Shaw. He returned to Australia inspired by Fabian ideas and the possibility of constructing a more rational society. Over time, his economics shifted towards classical liberalism, while his concern with social responsibility endured.

That evolution is admirable. Changing one’s mind is a valuable quality in a social scientist. It is indispensable in a democracy too.

Shann brought the same disposition to economic events. In The Boom of 1890 – and Now, published in 1927, he used the experience of the 1880s boom and subsequent depression to examine the risks Australia faced in the 1920s (Shann 1927). During the Depression itself, he moved between academic work and policy advice, contributing to debates over the Premiers’ Plan, unemployment, exchange‑rate flexibility and trade.

One colleague remembered Shann as ‘eager for the truth’. It’s a terrific phrase. In a learning state, policymakers should hold programs lightly – always curious for the truth, and always willing to adopt better approaches if the evidence points in a different direction.

Shann changed his views as his reading and experience changed. The institutional question for government is how to make that kind of intellectual movement easier.

How do we build a state that is, in Shann’s words, eager for the truth?

4. Australia is already running the experiments

The employment‑services trial that I began with is part of a broader movement. Across the Australian Government, randomised trials are increasingly being used to answer questions once settled largely through instinct and before‑and‑after comparisons.

Consider a Centrelink phone call.

Services Australia often sends a text message before an officer calls a customer. An initial analysis suggested that the text increased the chance of successful contact by around 5 percentage points. Other factors could have produced some of that difference.

So Services Australia worked with the Australian Centre for Evaluation to run a randomised trial. Across 23 teams, involving 275 service officers and nearly 5,000 outbound contacts, some customers received the existing text message, some received revised wording, and some received no advance message.

The trial produced a bigger estimate than the earlier observational analysis. A pre‑call text increased successful contact from 52 per cent to 64 per cent. Changing the wording brought similar benefits (Services Australia 2026).

A text message before a phone call is not the frontier of technological innovation. It is unlikely to feature in the next season of Black Mirror (which is probably a good thing). Yet a 12 percentage point improvement in successful contact, applied across a large service system, saves thousands of hours for customers and staff, and potentially millions of dollars for taxpayers.

It also shows how randomisation allows us to get closer to the truth. The earlier data pointed in the right direction, while substantially understating the effect. Observational studies can exaggerate causal effects or understate them. Random assignment helps distinguish the effect of the intervention from the circumstances surrounding it.

A series of experiments conducted by the Behavioural Economics Team of the Australian Government, BETA, with Treasury provides another example. The question was how to encourage bank customers to seek a better deal on their home loan or savings account.

In survey experiments, prompts had a significant impact. When people were shown a detailed home loan prompt comparing their current interest rate with the average rate on comparable loans, the share recommending some form of market engagement rose by 49 percentage points. A shorter prompt raised it by 16 percentage points (BETA 2026a).

Then BETA moved into the field.

The first field experiment sent home loan customers a fairly generic prompt. It sat below the main menu in the banking app, where customers had to scroll to find it. The contact rate was 2.9 per cent among people receiving the prompt and 2.7 per cent among the control group. The estimated effect was close to zero.

The second trial changed the design. The message was prominent and personalised, telling customers that they might be eligible for a lower rate on their particular loan. It was delivered through multiple channels. Contact with the bank almost doubled, from 4.8 per cent to 7.2 per cent (BETA 2026a).

The world is more complicated than ‘prompts work’ versus ‘prompts fail’. The first field trial showed that a message easily overlooked had little impact on behaviour. The second showed that prominence and personalisation recovered some of the effect in the real world.

That gap between a promising idea and its performance in the field is relevant to policy. Economists sometimes call it voltage drop: effects shrink as interventions move from controlled settings into everyday life (List 2022). Experiments can show us where the voltage is being lost.

Another Australian trial concerned the annual reporting obligations of charities.

Each year, registered charities need to submit an Annual Information statement to the Australian Charities and Not‑for‑profits Commission. The Australian Centre for Evaluation tested an additional reminder sent directly to a charity’s responsible person. Around 15,000 charities that had yet to lodge were randomly allocated between the intervention and control groups.

The extra reminder raised on‑time submission from 56 per cent to 62 per cent, with charities also lodging around 3 days earlier (Nguyen et al. 2025).

Much of government consists of millions of interactions involving forms, letters, websites and deadlines. A small improvement repeated across a large administrative system can yield a sizeable aggregate return.

Experimentation can also be used in the policy design process.

BETA recently examined how the design of the Paid Parental Leave claim process might affect the sharing of leave between parents. In a randomised survey experiment involving more than 1,500 participants, researchers changed the way sharing was presented in a hypothetical claim process. Making sharing more salient, alongside an explanation of the purpose of the scheme, led men to claim more leave in the exercise. The corresponding effect among women was small (BETA 2026b).

A survey experiment provides evidence about behaviour in a constructed setting. Its advantage is speed. Forms and interfaces can be tested while they remain cheap to change. A field trial can follow when the question warrants it.

Some Australian experiments tackle broader economic questions. Andrea La Nauze and her co‑authors studied whether electric‑vehicle owners would shift charging towards times when renewable electricity was plentiful (La Nauze et al. 2026).

They recruited 390 Australian Tesla owners and tracked their vehicles using minute‑by‑minute telematics. Participants were offered financial incentives to move charging towards the middle of the day or away from the evening peak.

Among owners without rooftop solar, midday charging rose by 34 per cent. Incentives to avoid the evening peak reduced charging during that period by 27 per cent among owners without solar and 17 per cent among those with solar.

The detailed, minute‑by‑minute data showed how people adjusted by tracking all driving, charging, and vehicle locations. Much of the extra midday charging occurred away from home, particularly among commuters. Evening reductions occurred mainly at home. Access to workplace charging and faster home chargers shaped how readily people could respond. The study illustrates the power of combining randomisation with big data.

A well‑conducted randomised trial can also keep producing information long after the original intervention ends. Take for example the right@home trial, begun in 2013. Researchers recruited 722 pregnant women experiencing adversity in Victoria and Tasmania. Half were offered a sustained program of around 2 dozen nurse home visits from pregnancy until their child turned 2; the other half were assigned to a control group, receiving the usual universal child and family health services (Goldfeld et al. 2017).

This was a trial of the program now known as Maternal Early Childhood Sustained Home‑visiting. By using a rigorous randomised trial, it was possible to quantify the program’s impact. Experts might have anticipated that nurse home visits would be beneficial, but nurse home visits are also relatively expensive. Only a rigorous evaluation can determine whether home visits are the most cost‑effective intervention.

As years went by, researchers continued following the families. By ages 4 and 5, the trial was yielding evidence of benefits for the mental health of children and mothers, but no evidence of benefits for children’s learning (Goldfeld et al. 2022). More recent work has linked the original trial to administrative data at school transition, extending the period over which its effects can be studied (Price et al. 2024). Researchers continue to assess the program’s cost‑effectiveness as its participants make their way through school.

As with other long‑running randomised trials, such as the 1962 Perry Preschool Program, the 1977 Diabetes Study, the 1991 Women’s Health Initiative and the 1994 Moving to Opportunity Study, the expensive step is creating a credible treatment and control group at the outset. Once that exists, the original randomisation can become a durable research asset. With consent, appropriate governance and good data linkage, outcomes can be examined years later.

A trial of an early‑childhood program can eventually tell us something about school readiness. In time, it may tell us about crime, employment, earnings or family formation. The randomisation occurred once; the knowledge can continue accumulating. To adapt an old saying: a society grows great when researchers start randomised trials whose results they know they may never publish.

These Australian experiments span a text message before a Centrelink call through to the timing of electric‑vehicle charging. Some produce big effects. Others provide insights about implementation. All of them increased our knowledge about what works.

5. From isolated experiments to an Australian learning system

Individual trials can improve individual programs. Even better is a system in which evaluation becomes an integral part of the way government operates.

Australia is already moving that way. The 2026 state of Evaluation report identified 847 evaluations underway, planned or recently completed across 41 Commonwealth agencies. Among 219 impact evaluations were 14 randomised trials conducted across 8 agencies. Thirty‑four of the 77 agencies responding to questions about evaluation capability had a dedicated evaluation unit, compared with around 18 in 2021. Another 5 planned to establish one by 2028 (Australian Centre for Evaluation 2026).

Admittedly, we are yet to conduct a randomised trial on the effectiveness of randomised trials. Yet those numbers suggest that high‑quality evaluations are becoming more common. And they provide room for improvement. Fewer than half of responding agencies had a senior executive responsible for evaluation, while only one‑third routinely published evaluation findings.

The Strengthening Evaluation in the Australian Government Action Plan 2026-2030 seeks to raise the bar (Australian Government 2026). Its 14 actions cover leadership, capability, planning and the use of evidence.

Government entities are encouraged to appoint chief evaluation officers or senior evaluation champions, and to establish evaluation committees or incorporate evaluation into existing governance. Large entities are also expected to maintain dedicated in‑house evaluation capacity. The plan also calls for entity‑level evaluation strategies that decide in advance which programs most warrant scrutiny.

From 2027, the Australian Centre for Evaluation and central agencies will consider stronger requirements for evaluation, including the possibility of mandatory evaluation for high‑value or high‑risk programs. The goal for 2030 is a proportionate evaluation system in which all such programs face formal evaluation and, wherever possible, impact evaluations explain how they established a credible counterfactual.

We also need to consider what happens after an evaluation is finished. The Action Plan calls for findings to be shared and published where appropriate, for formal management responses, and for evaluation evidence to feed into funding and budget processes. It’s about building rigorous evaluation into the culture and systems of government.

And we’re investing in evaluation skills and capacity. The Australian Centre for Evaluation provides specialist expertise at the centre of government. BETA brings behavioural science and experimental methods to policy questions. The APS Evaluation Profession is building a community of practitioners across agencies. Services Australia’s work with the Australian Centre for Evaluation led to its own Impact Evaluation Hub, giving staff guidance and tools to conduct further evaluations (Services Australia 2026).

Here in Western Australia, the Australian Centre for Student Equity and Success at Curtin University describes itself as a ‘What Works Centre’ for student equity. Its trials registry contains half-a-dozen studies using random assignment, testing interventions from bridging courses to support for disengaged students.

It helps that the data are far better than ever before. Australia’s linked administrative datasets have massively reduced the cost of following outcomes over time. The employment services experiment I began with could track employment, earnings and income support receipt through existing administrative records. The experiment would have been more expensive, and the sample size would have been smaller, if the data were collected in a survey.

Linked microdata changes the economics of experimentation, by reducing the cost of measurement. It can also extend the life of an experiment. As the right@home study illustrates, treatment assignment made years earlier can continue producing evidence when later administrative outcomes become available. Datasets such as Gen V, which covers 50,000 children born in Victoria, have been collected with explicit consent for cohort members to be invited to participate in randomised trials over their lifetimes.

A rigorous evaluation isn’t like a rigorous audit. Evaluation is far more useful when it is designed alongside a policy. Randomisation can be built into rollout, data requirements settled in advance and outcomes specified before anyone knows the result. Retrofitting a credible counterfactual years later is often much harder.

Philanthropic foundations can also help to create a learning society. Foundations can operate as the venture capital arm of government – engaging with experts, working with communities, and trialling promising ideas to figure out what works (and what does not).

In the US, Arnold Ventures, a foundation with a major program of funding randomised evaluations, supports replication, longer‑term follow‑up and the use of administrative data, while asking researchers to set out clearly how their work could inform policy (Arnold Ventures 2026). In the United Kingdom, the government is exploring how best to partner with foundations on randomised trials, exploring approaches in which government and philanthropy co‑fund rigorous evaluations, with government enabling the policy experiment and ensuring data access. At the early stage, such programs might have foundations as the majority funder, with the government later becoming the majority funder if the program proves successful.

In Australia, when the Paul Ramsay Foundation ran a $2.1 million Experimental Evaluation Open Grant Round in 2024, it received more than 100 applications for the 7 grants on offer – illustrating the enthusiasm among Australian nonprofits for running randomised trials.

This isn’t just about a handful of trials. It is about building systems that are continually testing, experimenting, adapting and improving.

6. What other countries are institutionalising

Australia can also learn from countries that are pushing experimentation further into the machinery of government.

Britain has taken this a long way. Its Evaluation Task Force sits across the Cabinet Office and Treasury, giving it a foothold in policy development and spending decisions. By September 2025 it had provided evaluation advice on more than 750 projects worth over £560 billion. A central Evaluation Registry, launched in 2025, now contains more than 2,000 evaluations (Evaluation Task Force 2026). Four‑fifths of English schools have taken part in an Education Endowment Foundation evaluation (Education Endowment Foundation 2025).

The broader British approach combines central expectations with resources that help departments evaluate well. A £22 million Evaluation Accelerator Fund has supported high‑quality evaluations of more than 50 priority projects. The Evaluation Task Force has also expanded evaluation training across the civil service (Evaluation Task Force 2026).

In 2026, Britain revised the Magenta Book, its central guide to evaluation, to give ‘Test and Learn’ a prominent place in policy development. The approach starts with the assumptions underlying a policy. Which are uncertain? Which can be tested quickly? Teams can begin at small scale and observe how the intervention operates. They can then adjust the design before moving towards rigorous impact evaluation as the policy develops (HM Treasury and Evaluation Task Force 2026).

Under the old model, evaluation often follows design and implementation. Test and Learn brings evidence generation forward, while important features of the policy remain open to adjustment.

Canada offers an institutional change that Australia could readily borrow.

When Canadian officials prepare submissions for the Treasury Board, the guidance asks them to identify gaps in the evidence supporting the proposed design. It then asks whether those gaps will be addressed deliberately through a pilot, a randomised trial or another form of experimentation (Treasury Board of Canada Secretariat 2026).

A government seeking approval to spend money is asked to specify what it knows about the proposal and where the uncertainties lie. It must then consider how those uncertainties might be reduced.

Australia has already begun strengthening evaluation capability. We have already borrowed from the UK Evaluation Taskforce’s training model to start training other public servants to be ‘evaluation trainers’ within their agencies.

The Canadian example suggests that when a major proposal comes forward, we might routinely ask, ‘What are we uncertain about, and how will we learn?’

The British example suggests another handy question: ‘What can we test before we scale?’

These are the kinds of questions that people in a learning state should be asking.

7. Federalism: Eight laboratories, if we choose to use them

Australia already possesses a structure that could generate more policy knowledge than we currently extract from it: the federation.

States and territories make different choices in schooling, hospitals, transport, planning and regulation. Economists often analyse those differences through the lens of fiscal federalism: decentralisation can allow policies to reflect local preferences and conditions, while creating scope for innovation (Oates 1999).

Occasionally, a state innovation spreads. Western Australia established Keystart in 1989, pioneering low‑deposit lending and later shared equity for households struggling to enter home ownership. Shared‑equity schemes subsequently appeared in other states, before the Commonwealth adopted the approach nationally through Help to Buy (Commonwealth of Australia 2024). Victoria led Australia on compulsory seatbelts in 1970 and random breath testing in 1976, both of which spread nationally (Australian Institute of Health and Welfare 2016). South Australia introduced container deposits in 1977, decades before every Australian jurisdiction had a scheme (Environment Protection Authority 2023).

Such examples are rarer than they should be. State policy innovation has generally occurred without evaluation designs capable of telling other jurisdictions, with much confidence, whether the policy caused the outcomes that followed. In the absence of strong evidence, whether a policy spreads across states can depend more on politics than on how well the policy works. Federalism could generate far more useful evidence if evaluation were built into policy variation from the start.

Eight different policies become 8 experiments only when the variation is structured for learning.

In other countries, federal governments use grants to encourage experimentation by states. For example, in the United States, the Second Chance Act, which funds strategies to help prisoners re‑enter the community, set aside 2 per cent of program funds for evaluations that ‘include, to the maximum extent possible, random assignment . . . and generate evidence on which re‑entry approaches and strategies are most effective’ (United States Congress 2008).

This approach recognises that rigorous evaluation by states is a public good, whose benefits spillover to other jurisdictions. That’s why the US federal government chose to subsidise states that conducted randomised trials – because all jurisdictions gain when we learn what works.

Australian federalism is sometimes criticised when it leads to duplication. But for a learning state, duplication could become experimental capacity. If we are going to have 8 systems, we may as well learn 8 times as much. The opportunity is to turn policy variation that already exists into evidence that every jurisdiction can use.

8. When the reasons run out: the case for lotteries

In some settings, a lottery is the fairest way for a government to allocate a scarce opportunity. Once applicants have crossed a reasonable threshold, further ranking can create a spurious precision. As QUT statistician Adrian Barnett puts it, the case for lotteries arises ‘when the reasons run out’ (Barnett 2025).

Australia has used lotteries for decisions carrying major consequences. Between 1965 and 1972, more than 800,000 men registered for national service. Around 63,000 were conscripted, with selection determined by a ballot of birth dates; more than 19,000 conscripts served in Vietnam (National Archives of Australia n.d.). A lottery made the allocation rule impartial.

Today, Australia uses random selection for a happier purpose. Under the Pacific Engagement Visa, eligible people from participating Pacific countries and Timor‑Leste enter an electronic ballot. Those selected can then apply for permanent residence, subject to meeting the visa criteria (Department of Home Affairs 2026).

In New Zealand, the Health Research Council has used a lottery for research grants. Reviewers first ask whether a proposal is potentially transformative and whether it is exploratory but viable. Proposals passing that threshold enter a lottery. In a survey, 63 per cent of applicants supported the system where weaker applications had first been screened out and the remaining proposals were roughly equal in merit (Barnett 2025).

Lotteries have also been used in medical education. The Netherlands used weighted lotteries for medical‑school admission for decades, giving applicants with stronger school results better odds while retaining an element of chance (10 Cate 2021). Sweden has used lotteries when medical‑school applicants were tied at the maximum high‑school grade (Chen, Persson and Polyakova 2022; Artmann et al. 2022).

Once such lotteries exist, researchers have something close to a randomised trial. Peter Siminski and his co‑authors have used Australia’s Vietnam ballots to study outcomes decades later. They found little evidence that army service increased long‑run mortality (Siminski and Ville 2011), but sizeable effects on later employment: among men induced to serve in Vietnam, employment in 2006 was 37 percentage points lower, alongside increased receipt of veterans’ disability pensions (Siminski 2013). Later work found adverse effects on hearing and mental health, while finding little effect on mortality, cancer or emergency hospital use (Johnston, Shields and Siminski 2016).

The medical‑school lotteries have produced the same research dividend. Using Dutch admissions lotteries, researchers found that the apparent health advantages enjoyed by parents of doctors largely disappeared under causal analysis: having a child become a doctor had little effect on longevity or most measures of healthcare use (Artmann et al. 2022). By contrast, a Swedish study using admissions lotteries found that doctors’ relatives were substantially more likely to engage in healthy behaviours (Chen, Persson and Polyakova 2022).

Lotteries can yield 2 benefits. They can allocate scarce opportunities when fine distinctions among qualified candidates are hard to defend. And they operate as randomised trials, allowing society to learn what works and helping us produce better policies.

9. The productivity of learning

In the early 1970s, the psychologist Donald Campbell imagined an ‘experimenting society’: one that would try out reforms, evaluate their consequences and move on when better alternatives emerged (later written up as Campbell 1991). His aim was what he called an ‘evolutionary, learning society’, where implementation itself helped generate knowledge. In other words, government should behave more like a good scientist and a less like someone defending a restaurant choice they made 20 years ago.

A learning society is also a more productive one. The Productivity Commission has argued for rigorous evaluation to improve outcomes and value for money from public spending; CEDA has called for evaluation to be built into program design; and the OECD has stressed the role of impact evaluation in improving resource allocation and public services (Productivity Commission 2017; Winzar, Tofts‑Len and Corpuz 2023; Jacobzone et al. 2025). CEDA’s review of community services found that weak evaluation could leave ineffective programs running while effective ones were stopped (Winzar, Tofts‑Len and Corpuz 2023).

There is also a democratic benefit. With trust in government under pressure, randomised trials offer a form of public humility: government can say that it has a view, while remaining willing to test it and change course. Unlike complex econometric identification techniques, a randomised trial is simple and straightforward. Policies gain legitimacy when governments can clearly explain the evidence that underpins them (Tanasoca and Leigh 2024; Leigh 2025).

In running randomised trials, a modern learning state can use linked data to reduce the cost of finding answers, and federalism to generate useful variation. It can deploy lotteries when there is no fairer way to allocate a scarce service and encourage philanthropic foundations to finance exploration. A learning state uses today’s policy rollout to improve tomorrow’s policy development process.

Shann was remembered as ‘introspective and self‑questioning’ (Snooks 1988). They’re important qualities in an economist, and pretty useful for public servants and politicians too. A government will make better choices over time if it is confident enough to test its ideas, and curious enough to learn from the results.

Australia already has what we need to make this happen: capable public servants, expanding data assets, a federation full of variation, and growing experience with rigorous trials. The challenge now is to make learning part of how government works. If we can do that, we will leave future governments something immensely valuable: a state that keeps getting better because it keeps finding out.

Acknowledgements

My thanks to Blair Arnold, Adrian Barnett, Sharon Goldfeld, David Halpern, Peter Siminski, Ana Tanasoca, Cassandra Winzar and officials in the Australian Centre for Evaluation for advice, suggestions and feedback on earlier drafts.

References

Arnold Ventures 2026, Strengthening Evidence: Support for RCTs to Evaluate Social Programs and Policies, request for proposals, Arnold Ventures.

Artmann, E., Oosterbeek, H. and van der Klaauw, B. 2022, ‘Do Doctors Improve the Health Care of Their Parents? Evidence from Admission Lotteries’, American Economic Journal: Applied Economics, vol. 14, no. 3, pp. 164-184.

Australian Centre for Evaluation 2026, State of Evaluation in the Australian Government 2026, Department of the Treasury, Canberra.

Australian Government 2026, Strengthening Evaluation in the Australian Government – Action Plan 2026-2030, Department of the Treasury, Canberra.

Australian Institute of Health and Welfare 2016, Australia’s Health 2016, AIHW, Canberra.

Barnett, A. 2025, Using Lotteries to Increase Fairness and Efficiency in Research Assessment, presentation to the DORA Asia‑Pacific Funder Discussion Group, 13 March.

Behavioural Economics Team of the Australian Government (BETA) 2026a, Prompting for a Better Deal: Understanding and Overcoming Barriers to Financial Product Switching, Department of the Prime Minister and Cabinet, Canberra.

Behavioural Economics Team of the Australian Government (BETA) 2026b, Paid Parental Leave Decisions, Department of the Prime Minister and Cabinet, Canberra.

Campbell, D.T. 1991, ‘Methods for the Experimenting Society’, Evaluation Practice, vol. 12, no. 3, pp. 223-260.

Chen, Y., Persson, P. and Polyakova, M. 2022, ‘The Roots of Health Inequality and the Value of Intrafamily Expertise’, American Economic Journal: Applied Economics, vol. 14, no. 3, pp. 185-223.

Commonwealth of Australia 2024, Help to Buy Act 2024 (Cth), No. 124.

Department of Employment and Workplace Relations, forthcoming, Reducing Mutual Obligations: A Randomised Controlled Trial in Online Employment Services, Department of Employment and Workplace Relations, Canberra.

/Public Release. View in full here.