Host: Let's welcome Matthew to the stage.

Title slide: AI to Support Educators — Esperanza/Ed Tech Month Hong Kong, October 13, 2023

Matthew Rascoff: Thank you so much. It is really a pleasure to be here. This is my first time in Hong Kong, which I'm a little bit embarrassed to say, but it's been a long time coming, and I have so admired the Asian educational cultures for so long. There's so much curiosity in the world that I come from about the story of the educational success of the systems here, and to get to experience it firsthand — with some of the expertise of faculty from the universities, but also the edtech innovators who are bringing their ideas to us — is truly inspiring to me.

I think one of the challenges of an event like this is that we all come from such different backgrounds and different contexts, and education systems are still very nationally rooted, despite the global efforts of organizations to diffuse ideas and innovations around the world, and despite great edtech innovators who are cosmopolitans crossing borders and spreading these ideas. So one preface that I would offer is that I come from the US system. I was trained in the US system. I've lived in other countries, I've studied in other countries, but I think there are some limitations, and what might be innovative in one country might be humdrum for you in another context. So bear with me on that: this comes from a US perspective.

When I teach at Stanford in the business school, I have 50 students in my class and 40 different countries represented, and I'm teaching them about the American education system. I always feel a little bit uneasy about the provincialism of a curriculum that's organized around learning about what's happening in education. But I console myself that they are the enablers of learning. They're going to go back, most of them, to their home countries, and they're going to spread ideas. There is a kind of cross-fertilization that happens through that process, in which there's intellectual arbitrage and new ideas that come from the exchange of ideas in context, just like this. So that's my caveat on the American perspective that I'm going to bring — and my gratitude to all of you for hosting me here and for listening in as I share what's happening at Stanford.

The title of this is "Empowering Students with AI," and the message that I want to leave you with is that the way to do that is to empower educators with AI — and that humans have a fundamental role that we cannot discount. Rather, we should be investing in it much more deeply, even as the value of human educators is being questioned.

I was listening very carefully this morning, and I did not hear a single time the term "personalized learning." If we were doing this conference five years ago, personalized learning would have been the beginning, middle, and end of this, but it has not come up a single time. Why is that? I'm asking genuinely. Why did we not talk about personalized learning today? Has anybody worked in this space long enough to know that fad — the rise of that fad and then the decline of that fad? I see some nods.

Why did personalized learning fail at the Chan Zuckerberg Foundation, such that they've just written off a $100 million investment in their Summit Learning platform, which was supposed to be the scalable mechanism for kids to do personalized learning in K–12 schools across the United States? They laid off all of their staff who were working on this. They've given up on this idea. So to me, this should make us somewhat skeptical of the kind of rhetoric that we heard from Ben and from George this morning. There are some fads and some faddish behaviors in this system that seek to chase the latest idea and sometimes lose track of the fundamentals — fundamentals that do not rise and fall on a five-year cycle but are much more oriented toward the long term. To me, that has to be our focus.

Chalkbeat headline: Mark Zuckerberg tried to revolutionize American education with technology. It didn't go as planned.

We were talking a little bit this morning about the patience that is required — John, this came up with you — the patience that's required to invest in this space. The investors don't necessarily get that; they're impatient, they want returns. But if you're an educator in this space, think about the deep impact for the long term that educators have had on you, individual teachers. Those are not fads. So to me, Zuckerberg and Chan Zuckerberg and Bill Gates talking about personalized learning — this was the Gates Foundation's core investment thesis — they've basically given up on it. And Salman Khan has talked about it too: personalized learning allows students to progress through content at their own pace without worrying about being too far behind or too far ahead of their classmates.

Where has that gotten us? Kids on computers in classrooms with headphones on, who are not learning with one another, who are not being socialized, who are not being helped to create an identity, who are not building a learning community. They are not progressing. And the data finally caught up. I credit Chan Zuckerberg for at least being honest about the lack of results and being willing to write off a $100 million investment.

So to me, the core fundamental does not change over time, no matter how advanced the technology is that you're building: we need to be investing in great educators and great teachers. And they actually do personalized learning. Great teachers are listening to a student's needs, and they are doing it systematically as part of what they do.

Humans make learning personal

This is an example from Dan Meyer, an educator who I love and who I highly recommend. He's a math educator who writes about math pedagogy, mostly in K–12, but I think a lot of these lessons are relevant to higher education as well. He has basically argued that an educator like Liz Clark Garvey, in New York City public schools, can start the lesson with a whole-class move — she'll ask one question for the whole class — and then, through moving around the class, listening to what the students say and how they decipher the problem, she is able to understand where students are and to meet their needs.

The challenge with a context like this is that it seems to depend on heroic individual teachers like Liz. There has not been a systematic mechanism — maybe in Singapore there is, maybe in Hong Kong there is, but in the US there has not been a systematic mechanism — to take a model like this from individual great educators and scale it to the order of the millions of teachers that we have. Three million teachers in our schools, not to mention higher education. So to me, the challenge is not how do we give every kid a laptop and a screen and headphones; the challenge is how do we give them a great educator who cares about them, who will create a learning community in a classroom of people who will learn together and support one another. That, to me, is the precious thing and the rare thing — the thing that has become even more precious and more rare under the conditions of technology seeming to take away some of the role for humanity in our classrooms.

So to me, if you care about personalized learning, or at least the ideas behind personalized learning, the way to do that is not to take the persons out; it's to bring the persons in, bring the humans in. But how do you do that systematically? How do you train an educator workforce to really capitalize on what makes us most human? As we were talking about earlier today, that is a non-trivial challenge — for teacher training programs, for in-service teacher training for teachers who are already in the workforce, and for faculty at universities, who sometimes get the least training in pedagogy, including at places like Stanford. That, to me, is how we should frame the challenge, and how we frame the challenge is going to determine what kinds of solutions rise to the top.

What I want to talk to you about today is a model that I see emerging at Stanford for how we can simultaneously democratize and humanize education. Usually these are a trade-off. Think about how MOOCs did a great job of democratizing education but took every last iota of humanity out of a course and totally denatured it. They basically turned a course into YouTube, for all intents and purposes. Even forums have been taken out of Coursera — why? Because it doesn't work when you have an always-on platform. They eliminated the idea of cohorts, which were there at the beginning; that's gone. So MOOCs did a great job of democratizing but a horrible job of humanizing education. They made it this cold, bloodless place to learn that most people cannot get through — only 6% of learners can get through a MOOC, for that reason.

But places like the liberal arts colleges that George was talking about this morning do a great job of humanizing education. I teach at Stanford; this quarter I'm teaching 14 students with two instructors. Fourteen students. Think about that ratio. Think about the luxury of having two instructors for 14 students. That is an incredibly wonderful privilege that is very expensive to provide. So we figured out how to humanize, but we have not figured out how to democratize, because the model of Stanford is predicated on extreme selectivity — and we want to be able to provide that to many more. But what is the mechanism by which you can do that? To me, this is a fundamental balance that we need to strike.

The model that I want to share comes from a project called Code in Place. This comes from a colleague of mine, Chris Piech, and his colleague Mehran Sahami, in the computer science department at Stanford. Has anybody heard of this project before? I'm so glad you haven't, because it means I'm sharing something new with you. It was launched during the pandemic, and it's now run in three separate cohorts and has reached 30,000 students. It has a unique pedagogical model that, to me, exemplifies what's possible in the balance of democratizing and humanizing education at the same time.

Democratizing and humanizing CS education at scale

This is the structure of Code in Place. I call it the "meso scale." The meso scale is the middle ground between the macro scale — the mega scale of a MOOC, which is on the order of millions — and the extreme micro scale of face-to-face learning at places like Stanford, which is on the order of tens. In between those is the meso scale. The way this model works is that it's a combination of synchronous and asynchronous learning, but all students are in a group of 10, and that group of 10 has a human section leader — a volunteer who's been trained by Stanford. Many of them are Stanford alums, though not exclusively. They are responsible for creating a learning community among those 10 students, and we had them in every single time zone around the world. This chart shows the geographic representation of our first cohort, and it basically shows you how we built this fractal model of scale, where every group of 10 is led by a section leader, and then every section leader themselves has a community, and those are led by training leaders.

The “meso” scale: students meet weekly in groups of 10 led by a section leader; section leaders are trained in cohorts of 20

Code in Place global reach

So what you end up with is 30,000 students — but it's not 30,000 students treated as one bulk of humanity. It's 30,000 students, each of whom is in a small section led by a caring individual who's been trained to support them: creating a cohort, creating accountability, creating community among that group of 10, in every single time zone in the world, in multiple languages. To teach 12,000 students in a cohort, we trained 1,200 co-instructors. It may be the largest co-taught course in history — that's what the faculty think. Twelve hundred co-instructors to reach 12,000 learners in a year, and that is the way to get to this kind of scale.

The question, though, is: what does that teaching look like, and how do you support 1,200 co-instructors? The results speak for themselves: 99% of our section leaders completed it, 56% of students completed — not bad for a free, non-credit course in which there is basically no skin in the game. Compare that to 6% for a MOOC. It scored 4.95 on Stanford's five-point teaching scale — take that for whatever it's worth. But the net promoter scores are an indication of whether somebody would recommend taking or teaching a course like this in the future: 30% of students said they would like to lead a section if it were offered again. So that's the result.

Code in Place impact: 99.6% of section leaders completed, 56% of students completed, net promoter scores of 90.3 and 70.1

How do you scale this, though? How do you get to 1,200 co-instructors? This, to me, is where the AI comes in. What I'm showing you here is an intervention that we designed — an AI feedback tool for the instructors — that ran as a randomized controlled trial behind the scenes in Code in Place. Some of the teaching fellows, some of the section leaders, got this intervention and others did not, and it had an N-size of about 1,100 co-instructors. This is the feedback that we were able to give those who received the experimental treatment. It may be a little bit hard to read, so I'll just call out some of the text that's in here: "This feedback gives you an opportunity to reflect and to support your professional development. It is not meant as an evaluation." So this is a formative tool to provide what's called instructional coaching.

Computers can help educators do what is uniquely human: AI-based feedback on your section

Instructional coaching is a proven method for improving teaching practices. It's a formative mechanism that brings an expert to the back of a classroom to give advice and feedback to a teacher. It's usually done in the context of K–12 education; it also happens in the higher education context. I ran a program like this in my previous role at Duke University. So it's a proven mechanism, but it's not scalable — it's very expensive, it's underprovided. It's effective, but it's not cost-effective. What we built here was an AI instructional coaching mechanism. The instructional coach in this AI was somewhat limited: it could only understand English — we were teaching this in some other languages, but it could only understand English — and it was basically measuring two mechanisms of student engagement: student talk time (how much do the students talk?) and moments when you built on student contributions.

Student talk time is an important concept to understand. This is a translation of research that was done by Rachel Lotan at Stanford — I see some nods from the psychologists here. It's critical for understanding classroom equity and classroom effectiveness, and it is predicated on the idea that if students are not talking, they're probably not learning. Rachel Lotan did this pioneering research in labs at Stanford that showed that student talk time, and equity of talk time, is critical not just for demonstrating their learning but for actual active learning to happen. So this is a measure of how much students are doing the talking in their class, and the quality of the uptake — the quality of the response that faculty, or the section leaders in this case, are giving to the students. Are they hearing what the students are saying and responding to it? "Thank you for that point, I agree with you, let me build on that point" — that kind of comment is a high-quality piece of feedback that's very hard to come by, and it's very hard to coach. The AI is there to do that in a non-judgmental way.

So there's no evaluation, there's no supervisor listening in. It's a little bit different than the classroom-observation model, where there's a kind of supervision component to it — this is not that. It's basically, I think, an effective use of the AI as a coach who's not there to judge you and just wants to help you get better. It turns out this mechanism is very effective for helping teachers improve.

What you can see here is a related model from a program called TeachFX, a Stanford student–created startup that's now a commercial edtech company and uses a similar mechanism. What I just showed you previously is called Empowered, which was the basis of the randomized controlled trial that we ran inside Code in Place. TeachFX is a commercial product that any school can purchase, and it shows how they measure the distribution of talk time in a classroom. What TeachFX has found is that just giving teachers a single recording with an analysis of how much they spoke and how much the students spoke is enough to significantly affect the second time they teach. There's a measurable impact on the quality of their teaching just from hearing this data once — just from seeing the data once, from their own classroom, they improve the next time. So it's a very effective mechanism for helping them improve.

How to scale human-centered teaching: the distribution of talk time in a lesson

TeachFX has recently built new AI tools that have more quality measures — so not just quantity of talk time (which we know is really important) and the distribution between the teacher and the students, but also more semantic analysis, using natural language processing, of what is actually happening. Similar to what you saw in the Code in Place mechanism of the uptake of a student idea, you can see the distribution of who's talking when. You can replay an audio clip, and in this moment, this is feedback that's being given to the instructor: "You are guiding a discussion about data usage and units" — the AI is understanding this and getting it from the transcript of a classroom, an online or a face-to-face classroom, either one — "building on the students' contributions, you prompt the class to find the connection between Ria's comment and the topic at hand, encouraging them to think critically. You also acknowledge a correct response from Deborah, further engaging the conversation." That was written by a machine that's listening in on their classroom and giving feedback on the richness of the discourse that's happening in the classroom.

How to scale human-centered teaching: AI feedback on building from students' contributions

This, to me, is the potential of AI. I don't care about ChatGPT — if every teacher had this, we would see immediate improvements in classrooms. This, to me, makes ChatGPT an absolute distraction. We know this is effective. This is backed by years of data, backed by educational research and cognitive science research, and every teacher in the world would benefit from having an instructional coach like this. And we have the randomized controlled trial data that comes out of the Code in Place project to demonstrate the effectiveness. This is a working paper that my colleague Dora Demszky, who led this research and is my critical research partner in this project, published with Chris Piech, who led Code in Place. It showed that instructors' uptake of student contributions improved 133%. This is gold-standard, randomized controlled trial research — I encourage you to check out this paper. That is a huge improvement in educational terms: a 133% improvement in this gold-standard mechanism for the increase in teacher uptake, meaning they're pulling an idea back from the student, repeating it, affirming it, and building on it — high-quality discourse in the classroom.

Randomized controlled trials demonstrate impact

So to me, this is the opportunity that we have in AI, and it is not actually about AI for students at all. I think chatbots are bordering on a complete waste of time, because their hallucinations are totally unreliable — they're worse than calculators. Calculators don't hallucinate. ChatGPT is not a reliable source of information right now. Somebody said this morning that there is no bias and it doesn't care whether you're black or white. It absolutely cares whether you're black or white, because the programming in ChatGPT is known to be biased. It is known to be biased. And for most students, writing is a form of thinking. If you're taking away that mechanism for doing the thinking, how do you think we're going to get more critical thinking out of students, if we basically deskill them systematically from doing the thinking that is embedded in writing? So I tell my students: you're cheating yourself if you use ChatGPT. There is no good detector for it — even OpenAI basically admits that — but I think it is a blind alley with respect to learning.

But natural language processing, based on the same large language models, has enormous potential as a mechanism for empowering teachers and for giving teachers, educators, and instructors the feedback they need — and I have the data to prove it. That, to me, feels like the conversation that we need to be having about artificial intelligence and the future of learning: how we can double down on the unique human qualities that teachers have, on the impact that they have on us, on the impact that our classmates have on us, and avoid the distraction and the faddishness that you saw in the personalized learning movement — which can easily happen with educational AI. The real risk there is that we throw out the whole thing and lose the enormous potential of technologies like this to improve learning systematically.

So with Dora and our colleagues, we're about to publish a white paper focused on how to empower educators with AI. What we're laying out is a set of principles and a set of possible strategic directions that are informed by educational research — that are not driven by commercial hype from AI companies that know nothing about teaching and should not be trusted with your student data, but are driven by what we know about learning, driven by the cognitive science, the computer science, the learning science. This is a group of faculty, not just at Stanford but from other institutions, that is going to put this white paper out, and I will circulate it to this community when it is published.

The basic idea is that we need a set of principles to start with and a set of directions that we can agree upon. This was the starting point that we came up with at a conference a few months ago. It is an initial starting point, and I hope in the discussion we can delve into what we might have missed here, because I don't think it's comprehensive yet. I don't think it's global, necessarily, and I think the AI considerations in authoritarian states may be different than in different kinds of states — so it's worth thinking about different political contexts and different educational contexts when it comes to this.

Emerging principles: begin with equity, center teacher needs, promote high-quality instruction, build and inform educational theory

Beginning with equity: it means that we need to be invested in AI that raises the floor, not just focuses on the ceiling — that we need to be thinking, like Rachel Lotan does in her research, about what the bias is in a classroom for who gets how much airtime, and how we use the AI to rebalance class time and classroom participation. That is not going to happen if we don't design for it, if we don't design to mitigate some of the biases that were talked about this morning. That, to me, is a critical educational goal that is not going to come from commercial AI; it has to be driven by educators. We have to demand it, and edtech companies need to build for it — a set of features, a set of goals around educational equity, the floor and not just the ceiling.

Centering teacher needs: I've already talked a little bit about that. To me it goes beyond just professional development. There's all sorts of busywork that teachers have to do that could effectively be supported by AI, and I am excited about the potential, for example, of AI to do machine grading when it's possible. There's a company called Gradescope — it was a student project at the University of California, Berkeley — that has been really effective in a human–AI combination that makes grading, not just of multiple-choice assessments but also of open-ended items, much more effective. That, to me, is a model of what's possible in this space. Can we relieve the burden on teachers so they can spend more time doing the higher-value work with their students — spending more time caring for their students, listening to their students, helping their students form an educational identity, promoting high-quality instruction?

Strategic directions: professional learning, routine teacher tasks, curriculum development, formative assessment, research–edtech partnerships

That's basically what I was talking about before. We know a lot about what high-quality instruction looks like, and it is incredibly rare to find. This, to me, feels like the professional development mandate that we all have: build the tools that scale high-quality instruction. Most of us get to experience it, if you're lucky, once in your lifetime. Close your eyes and picture the great teacher that you had. Picture the great classmate that you had who contributed to your learning. How do we get that available to more people? That, to me, seems to be the goal. John, you were talking about going to the Stuyvesant school and learning with a great classmate, Eric Holder — 50 years later, that person is still in your mind, the impact that they had on you. How do we give experiences like that to more students? That is an absolutely critical question for us, if we want to democratize high-quality education — to make teaching like that, and classmates like that, available to more people — and to build an informed educational theory.

So this is the role of a research university, and we have many research universities represented here. What's so exciting about a project like Code in Place is that it was simultaneously an educational outreach project — I lead the digital education office, so we're supporting this to democratize learning — but it was a research project on the back end. The students didn't know that; they didn't need to know that. They opted in, actually, to a research program — the IRB required it, it was supervised research. To me, finding those synergies that both advance opportunities to learn and advance the research on the back end is powerful. There was a massive data collection that happened in that project that could not otherwise have been done. So it was advantageous both for advancing the research agenda for Dora and her team in the education school at Stanford and for democratizing learning.

I hope we can develop a model for new kinds of research–practice partnerships that bring in edtech, where a lot of that data now lies. Think about it: instead of a researcher and a practitioner, it's actually more of a triangle, in which there's an edtech company that has the data, there's a school that's implementing some intervention, and then there's the researcher, and the three of them have this new kind of trilateral collaboration. RPIP is a new acronym that I see coming into form — research–practice–industry partnerships — and the three of them each might have a role to play. So we might end up with something that looks like this virtuous cycle of improvement. (It looks like we lost the text on the top right.) Cognitive science, like Rachel Lotan's research, gets translated into teaching and learning test beds, like Code in Place, implemented in data-driven tools like the Empowered and TeachFX tools that give feedback to educators to try to improve their performance — and then that informs the data science research, which faculty like Dora and Chris Piech depend on. They depend on high-quality data sets in order to get these analyses, in order to publish their papers and advance what we know about how people learn.

A virtuous cycle of improvement: cognitive science research, teaching and learning test-beds, data-driven tools, and data science research

So I'll close there. Thank you so much for your engagement, and I look forward to the questions that come in the next session.

Thank you