All Episodes

Listen in on Jane Street’s Ron Minsky as he has conversations with engineers working on everything from clock synchronization to reliable multicast, build systems to reconfigurable hardware. Get a peek at how Jane Street approaches problems, and how those ideas relate to tech more broadly.

Learning Goals and the Goals of Learning: Teaching in the Age of AI

with Aaron Bauer

Episode 30   |   September 23rd, 2026

BLURB

Aaron Bauer is a software engineer and one of Jane Street’s few developer educators—a role that splits his time between writing code and teaching other people how to write it. Before joining the firm, he taught computer science at Carleton College, including four straight terms online during the pandemic. In this episode, Ron and Aaron discuss what it takes to teach engineering inside a company with its own language, its own version control, and its own editors, and what changes when an LLM can do the exercise for you. Along the way, they consider the underrated power of a live lecture; why the text editor is still the software engineer’s home base; using editor telemetry to find out how AI is actually changing developer workflows; and how Jane Street rebuilt its intern curriculum around testing, design, and code review now that producing the code is the easy part.

SUMMARY

Aaron Bauer is a software engineer and one of Jane Street’s few developer educators—a role that splits his time between writing code and teaching other people how to write it. Before joining the firm, he taught computer science at Carleton College, including four straight terms online during the pandemic. In this episode, Ron and Aaron discuss what it takes to teach engineering inside a company with its own language, its own version control, and its own editors, and what changes when an LLM can do the exercise for you. Along the way, they consider the underrated power of a live lecture; why the text editor is still the software engineer’s home base; using editor telemetry to find out how AI is actually changing developer workflows; and how Jane Street rebuilt its intern curriculum around testing, design, and code review now that producing the code is the easy part.

Some links to topics that came up in the discussion:

CONTENTS

  • Asteroids, Python, and a teacher on a beach (00:00:02)
  • Grad school, rendering, and Foldit (00:04:05)
  • Carleton, COVID, and the flipped classroom (00:06:43)
  • In defense of the lecture (00:10:02)
  • Feedback loops and mastery learning (00:14:54)
  • Inventing the developer educator job (00:21:29)
  • More and more to teach (00:26:48)
  • The Editors team, and what a “feature” is (00:31:18)
  • The editor as home base (00:34:10)
  • Giant diffs and other problems of scale (00:37:47)
  • Python notebooks (00:40:41)
  • Build it ourselves, or push it upstream? (00:43:30)
  • What agentic AI changes for editors (00:48:05)
  • Claude Code’s learning mode and mental models (00:53:27)
  • Teaching when the LLM can do the exercise (00:58:20)
  • Bootcamps, teach-ins, and a course catalog (01:07:26)
  • What are the learning goals? (01:16:15)
  • Rebuilding the intern curriculum for AI (01:25:10)

TRANSCRIPT

Asteroids, Python, and a teacher on a beach (00:00:02)

Ron

All right. Well, it’s my great pleasure to have Aaron Bauer with me. Aaron is a software engineer and one of Jane Street’s proud but few developer educators. We’ll talk a little bit more about what that means in a bit, but thanks for joining me.

Aaron

I’m very excited to be here.

Ron

So, I like to always hear a little bit about how people started out. You’ve done a lot of work, which is kind of on the overlap of both software engineering and computer science and teaching. And I’m kind of curious, how’d you get there?

Aaron

Sure. So my origin story you might say is I was sort of fascinated with computers from a child, computer programming always felt kind of intimidating. And then I did this summer program before my senior year of high school, creatively named the Summer Science Program, where I went to New Mexico for six weeks and they taught us a bunch of math, astronomy, physics, and some Python. We took telescope observations of an asteroid and then wrote some Python code to take those measurements and compute an orbit for the asteroid. So really fun and interesting program, but I came out of it or went into it being like, all these things seem interesting. And I came out of it being like, wow, the programming was just way more fun and really my thing. So at that time, my small town high school didn’t offer AP computer science, but found a retired teacher on a beach in North Carolina who I could take it from.

And then sort of from there -

Ron

How did you find the retired teacher on the beach?

Aaron

So North Carolina had an online course program that you didn’t have to live in North Carolina to use.

Ron

Oh wow.

Aaron

And it just turned out the guy teaching this course was not currently working in a school, but instead was on the beach.

Ron

Oh, amazing.

Aaron

And then the teaching side comes in that before, the summer before I go into grad school, I go back to the summer science program as a TA. And that’s sort of my first exposure to writing an exercise for students to do and giving a little lecture on binary search trees. And I was hooked and sort of from then knew that I wanted to do something with teaching.

Ron

It’s funny, by the way, just the science and physics hook and programming, that’s actually a deep thing in the history of computer science. A lot of the early programming and programming languages and work was all about high-end physics simulations. Every now and then when I run into someone who’s gathering telescope data, just the hardcore computer science of weird FPGA exploits and stuff that people are doing there, I always find very impressive. There’s just a lot of interesting pushing of the state of the art and also a lot of messy not understanding a lot about computer science that I feel like comes out of that scientific computing world.

Aaron

Yes. Something I did in undergrad was work with a Fortran simulation of impacts, kind of asteroids or comets hitting the surface. And this came out of code that was used to simulate atomic explosions in the early days of computing because the physics of those two situations are not that different, it turns out.

Ron

Oh, wow. Yeah, and there’s a bunch of super cool stuff in the early Fortran work. We care a lot about automatic differentiation these days because we write all these machine learning models and you need gradient descent. And for that, it’s nice to have derivatives. But automatic differentiation goes back to the early Fortran days where there was a bunch of Fortran compilers that could do this. It was messy. They could do first derivatives, but not second derivatives. You would take the derivative of one program, you’d get another program. And if you tried to take the derivative of that second program, God help you. Nothing good would happen there.

Aaron

Yeah, it’s very much computer science is a means to an end. I’m not programming because I love programming. I’m using it just as a tool because it beats a slide rule, I suppose.

Ron

Right. Okay, but you loved programming.

Aaron

Yes, I certainly did.

Grad school, rendering, and Foldit (00:04:05)

Ron

And then you went to grad school. Why did you go to grad school and what did you study there?

Aaron

So I didn’t know I was going to grad school until I took a computer graphics course my last year of undergrad and got incredibly excited about rendering and was like, “This is something I could do for the rest of my life. What’s a way that you get to do one thing for the rest of your life? You go get a PhD in that thing.” So I applied to some programs, got into them, and then as I was talking to people, realized that people weren’t so much doing research in rendering anymore. A lot of the field had sort of moved on. It was very mature. The remaining work was largely in industry. So inertia carried me into grad school anyway and kind of did some game design research, some computer science education research, and then ended up writing a dissertation on what you might call human computer problem solving, studying how players of the scientific discovery game Foldit, solved problems and collaborated and sort of worked with the automated tools available there.

Ron

Oh right. Foldit was the early, here’s how we’re going to solve protein folding by having people go off and try and solve protein folding as a game.

Aaron

Exactly. Humans have some big picture geometrical insight of, “Oh, I sort of know how these amino acids tend to fold up and what if we took this part and flipped it over here?” And at that time, computers were really good at the, “I’m just going to wiggle the protein a lot to find the lowest energy kind of local optima.” And those two things together were able to resolve proteins that had gone unresolved in the lab for many years. People playing this game often had no biochemistry background. So it was very successful and eventually moved on to people actually designing novel proteins in the game. But I think AlphaFold and other things have sort of eaten their lunch.

Ron

Do you know if some of the data that went into AlphaFold and the like came out of Foldit?

Aaron

I assume AlphaFold was trained on all known structures, in which case, absolutely.

Ron

So there’s a little bit of an educational point in that, right? You’re talking about how do you get a bunch of people who do not know a domain to learn it and how does that process work and the gamifying the education process essentially?

Aaron

Yeah. A lot of my work in grad school had this sort of cognitive psychology element to it of trying to understand how people think or learn and then how they apply those skills.

Carleton, COVID, and the flipped classroom (00:06:43)

Ron

Okay. And then where’d you go from grad school?

Aaron

So I knew that I wanted to go somewhere where teaching was my primary responsibility. And so that’s how I ended up at a small liberal arts college, Carleton College in Minnesota. I went there and taught two quarters in person, and then it was spring 2020, and then I taught four quarters online.

Ron

So this is a hard introduction to teaching.

Aaron

Yes, it was challenging for me, very challenging for the students, but also involved a lot of trying new things and experimenting and different formats, which I think is an important part of succeeding in education.

Ron

So how well do you feel like Carleton did at adapting to these new online constraints? And just even forget Carleton, I guess, your own classes, how successful do you feel like you were actually able to be at this?

Aaron

I think I eventually got to the point where I was fairly successful. I started out taking the model of that online AP computer science course that I had taken in high school, which was entirely text-based. There were no video lectures. You just read some stuff and you did some exercises. So I started out on that in part due to guidance of people were at home, not everyone had good internet connections, want to make this very accessible. But I just immediately got feedback that this really does not work for all. It worked great for me, does not work for all students. So then -

Ron

Accessible in some way, deeply inaccessible in others.

Aaron

Yes. So I experimented with flipped classroom approaches with synchronous classes splitting out into small working groups at various points.

Ron

So there’s a bit of jargon that I bet not everyone understands. Can you say a little bit more about what a flipped classroom is?

Aaron

Yeah. So a flipped classroom is instead of the class time being when the professor would deliver a lecture, the professor prepares a lecture, a recorded video ahead of time. Students watch that before class, and then you use class time to do exercises or discussions, something more interactive.

Ron

Right. And this is not a new idea that came up in COVID. People have been talking about flipped classrooms for a long time.

Aaron

Yes. This is a longstanding idea in education that has been implemented in different ways, but it gained a sudden relevance and I think widespread adoption when classes had to move online.

Ron

Right. And I guess part of the point of the flipped classroom is the in-person time, being able to leverage the in-person time to do things that are better than just giving a lecture.

Aaron

Yes. So you take the time where students are say, passively receiving some content or seeing worked examples and you say, “That does not need to be done in person. And we’re going to leverage the instructor’s ability to tutor someone one-on-one or to lead a discussion,” things that can only happen when you have time together in the classroom. Where I personally ended up was just trying to make, not doing the flipped classroom, but trying to make the time we were together more of a mix of interactive things and some instruction.

In defense of the lecture (00:10:02)

Ron

Yeah. I’m curious how much you buy the kind of argument behind the flipped classroom. Because my instinct about this has always been that people underrate the power of lectures, that lectures are kind of a form of performance that is actually really compelling. And the really being in person with the person doing the lecture, even though in some sense it seems dumb and how could you possibly need it? And surely you should just record the best lecture that has been given on the topic and watch it. That actually with real human beings in real environments, this just kind of isn’t true. And there’s just something more compelling about being there in person. It’s not like I think the flipped classroom idea is totally wrong, but I do feel like the whole thing of, oh, there’s just the passive lecture is just missing something about the human psychology there.

Aaron

Yes, I would tend to say that lines up with my experience, that there is something about having a person, as you say, perform live in front of you that I think has a sort of focusing effect of your attention. And because it is live, I think you’re more likely to try and take notes, which I think is good for people understanding, which is not to say you couldn’t do that when you’re watching a video in your dorm room, but I think it’s less likely. I think a solid hour of someone just talking at you, that starts to lose the engagement and become counterproductive. So another bit of jargon, active learning is very important for these sort of in-person things where throughout the lecture, I want to stop and pose a question and have people talk to their neighbors or have people stop and do a little problem on their own to check their understanding, things like this.

Ron

By the way, note-taking is another great example of a super weird and also strangely effective learning technique where just being in the classroom and taking notes, and then if you want, you just light them on fire. You never look at them again. There’s actually, as I understand it, pretty good evidence that this actually helps retention in a serious way, even though the notes themselves are totally useless.

Aaron

Having to process things and then distill them down to something that you write down. And then I agree. The evidence seems to be that the act of writing it down helps set it in your memory in some qualitatively different way than just hearing it, even if you understood it.

Ron

Right. And apparently writing is better than typing for this. Maybe I’m getting…This is what I thought was true about the research. It’s not like I remember the study or anything.

Aaron

That sounds familiar to me. I would say my own experience in the classroom is it’s maybe less about the typing per se than someone having a screen with all the information and distraction that humanity has created. Sitting between them and the presenter tends to undermine one’s ability to pay attention.

Ron

That seems very believable.

Aaron

This is also why I have generally tried to avoid teaching in computer labs, which is something that people often do in CS courses. But I just can tell as soon as there’s a screen in front of people, 50% of people, even if they want to pay attention, they just, they can’t.

Ron

Yeah, exactly, this echoes a little bit my own experience, just teaching my kids how to program. One of the things I did a lot of was have them think about programs on paper. In fact, one of the things I conned my kids into thinking was an acceptable distraction when waiting for food to arrive at a restaurant was I would write a little Scheme program on a piece of paper with obfuscated names, and then they would have the puzzle of trying to figure out what did Flube do? And somehow I convinced them this was an acceptable way to spend time. But I do think the thing of talking to them about what a program means and starting out by being like, “I’m going to write down on the board or on a piece of paper, here’s an expression. What does this evaluate to?” I think the moment you give them a computer to type it in and try to evaluate it, they were playing the game of getting the computer to do thing and not thinking about the question of what is actually happening.

And I think there were ideas you could get across much more easily when you got the computer out of the way.

Aaron

Yeah. And there is definitely a thing that people just learning programming will tend to do, which is kind of arcane instructions for the machine. I will perturb it arbitrarily until it does what I want. And this is very easy to do when it’s on a screen and you can run it. But if I ask you to write out a program on paper and then decide when you think it’s correct, it’s more likely to force you to think about what it’s actually doing.

Feedback loops and mastery learning (00:14:54)

Ron

For sure. Okay. So you were at Carleton during this very complicated and disruptive COVID pandemic. You did a bunch of experiments and learned things. What happened when you got back to the classroom and people could be there? Are there things that you took away from that experience that you kept? What was persistent that you learned there?

Aaron

Yeah, I think one of the persistent things was I just found it very useful to have a lecture recording. I think this varies by institution. I’ve heard from other instructors that if they record the lecture, no one shows up to class because they can watch it. This wasn’t the case, Carleton students wanted to come to class and I tried to make being in class with this sort of interactive stuff actually useful.

Ron

It might be institution dependent, it also might be lecturer dependent.

Aaron

This is fair. This is fair. The classrooms at Carleton were not equipped generally with video equipment. So I was carrying around a webcam on a tripod to each of my classes and plugging into my laptop and recording the lectures and posting them online. And because I would get questions throughout the semester like, oh, I have this track meet or this obligation. I have to miss this class. What should I do? And if I have a recording, it’s like, well, it’s there. You can watch it. Or if someone asks a question, I can send them a link with the timestamp of the lecture. Here’s where I explained this thing. So that carried over.

I think I may have arrived at this even without COVID lockdown. But during that, I experimented a lot with automated feedback in various ways. So weekly quizzes that were multiple choice and auto-graded or different sorts of fill in the blank kind of things, different auto grading software where students could submit homework and get immediate feedback on whether it passed tests, that sort of thing. And carried that forward into my courses and found it because getting feedback on your work is an essential part of learning. And I was always dissatisfied with, okay, you do the homework, you turn it in. A week or two later, an instructor or TA has graded it and you give back some feedback, but at that point you’ve moved on and you may not even read the feedback that you get. Also, incorporated a technique that is sometimes called mastery learning, though I didn’t do it in its fullest extent, but the idea is that you have to demonstrate mastery or competence in a thing before you can move on, where you can attempt at any number of times and you keep getting feedback and you sort of re-attempt until you actually have gotten it right.

And so some of this automated assessment let me take more of that approach.

Ron

Interesting. My first instinct about the mastery thing is I would’ve done very poorly with that when I took math classes in college, because I felt like the experience of taking high-end math classes was like I was holding on, barely keeping up, trying to understand what happened. And then six months later, someone else was taking the class and asked me about it, be like, “Oh yeah, I know how that works.” I couldn’t quite get my arms around it at the time, but somehow just the passage of time after. And then you look back, you’re like, “Oh, yeah, I know how that fits together.” But I’m probably under-imagining, like probably I could have demonstrated some local mastery. I think the demonstration of mastery and also feeling really confused are probably like, consistent experiences.

Aaron

Yeah, I think this is often applied in a sort of maximal extent with much younger students. In elementary school math class, you have to like, it’s sort of self-paced in a way where the class might be at different points through the material and you have to master addition or the long division algorithm before you move on to the next unit. I think as you get higher level, you start needing to apply this in more targeted ways and that the flexibility is not there to the same extent in say a college course.

Ron

Did you actually do the thing of having people stay on older topics until they demonstrated mastery? Or was it more like there’s a thing that you should do, which is keep on iterating until you pass, but that happens at a fixed time and you have to do it at a particular point in time in the semester?

Aaron

Yeah, it would not have worked for people to progress through different rates, not least of which because students can procrastinate, it’s a thing people do. And I think it would’ve been a disservice to let people pile up all the work until the very end. So you do have to give people deadlines throughout for their own good and force people to move on. So it was a model of you have this short quiz to do this week. You can attempt it any number of times, but you are just going to get the best result of the times that you take it. And the questions are set up in a way such that it’s pretty annoying to brute force. The order of them keeps changing, the order of the answers keeps changing. So it’s just going to be faster if you actually try and understand all of it and then get a hundred percent that way.

Ron

Have you thought much about whether in the modern world some of this feedback, the kind of instant feedback that you think would be more useful for students can be generated by LLMs?

Aaron

Yes. I think there’s quite a lot of potential for that, particularly because when you have, say, a more structured thing to narrow the focus of the LLM. So for example, feedback on code style as opposed to an essay. I think I see a lot more promise in addition to some automated grading of test cases and maybe some linter to give students feedback of, like, you need to follow these rules. Some additional kind of code style pass I can see being very useful. But that being said, in my experience, there’s also plenty of false positives of LLMs giving feedback that is misleading. And so I think at the very least you would want to frame it carefully to students at least, or find ways to really reduce the rate at which it’s leading them astray.

Inventing the developer educator job (00:21:29)

Ron

Sure. Somehow you ended up here. How’d that happen?

Aaron

Right. So I was living and teaching in Minnesota, and my brother and his wife had moved to Brooklyn. She’s from Bay Ridge, and I decided that I wanted to live near them. My brother and I are close, so I decided to move to Brooklyn and was not going to be able to keep teaching in Minnesota while living in Brooklyn. And it was very fortunate that at that exact time, Jane Street was trying to hire for this weird educator programmer combo. I had just assumed, like, there were some things I liked about academia, but I was ready to try something new and didn’t really want to go through the academic job market. So I figured I would go be a software engineer somewhere. But it’s been amazing to continue doing education stuff in addition to doing some programming.

Ron

Yeah. And it’s maybe worth saying a little bit about the idea that we had in creating this job. We really wanted, I mean, to take a step back, teaching has been an important part of what we do at Jane Street for a very long time. In part, this is because we do some weird stuff that is not done so much in the outside world. Some of that is just subject matter stuff of nobody comes in knowing much about trading. And so there’s a big effort in teaching people to understand how to think about trading. And then also our tech stack is in a bunch of ways weird. And so there’s a lot of things we want to educate people about to understand the specifics of the programming language and the environment and the tools, a lot of which is very different from what you see on the outside.

And so we’ve spent a lot of energy on education for a long time now. And at some point we’re like, “It’s great to have people who are on the ground involved in things diving in and doing the educational work, but maybe we could get some people who have some expertise in the area and have actually done it for a material period of time and try and leverage that effectively. And so, well, let’s just actually create a job that’s just for this.” And it seemed clear to us that we wanted the people who were doing the teaching to also be people who were practitioners. Again, in part because there’s lots of things that are unique about our environment. And if the people involved don’t directly experience and work in the same ecosystem, they’re just not going to be effective or compelling teachers there. And so we designed this as a kind of 50/50 split where the idea is you’re going to spend half your time programming and half your time focusing on educational work, some of which is going to be directly teaching classes or creating curriculum, and some of which will be helping other people grow as teachers.

Because we actually don’t want there to be a small number of developer educators who do all of the teaching work, but to help grow the practice and make people better at it. I guess the other thing that I find interesting and a little frustrating about your description is that putting up this job ad did not at all work in the way I wanted it to, which I just sort of thought this was an obviously awesome job and we will advertise it and great candidates will fall from the sky because it’s a cool and unique job opportunity. And it turns out if we had put this out three months differently or something, we would’ve not gotten you at all. And in fact, actually as an ongoing thing, it’s actually been pretty hard to find the right people for this spot. I think it is less the case that people, in general, people don’t fall from the sky.

I think hiring people is really hard.

Aaron

Yes. And I think that this job is, as you say, very cool, but also very unusual. And the kind of person that would be looking for this job is also therefore kind of unusual. In particular, most people with a lot of teaching experience are in academia, and most people chose academia because they are passionate about what they get to do there and often are not looking to leave academia and work at a place in industry. The job is also a kind of unusual marriage of skills. If I’m teaching courses at a university, I’m not necessarily doing a lot of day-to-day software engineering, and I may or may not even be interested in that being a significant fraction of what I do.

Ron

Right. And we really want, at least in the ideal case, to find people who are really good educators who have a love for the craft and have spent a lot of time doing it and really think hard about what kind of things might you do and want to experiment in the space and learn how to be really better and are also just fully credible Jane Street software engineers. And that’s the intersection of two things where that intersection is relatively small. It’s maybe worth saying that we’re really excited with how having people in this role has turned out. There’s some amount of taking some internal people and converting them into this role, which is another path. But we’re also excited to get more people from the outside to do this. And I think there are just more and more domains where we want it. Maybe it’s worth talking through some of the other spots that we are trying to find developer educators for.

More and more to teach (00:26:48)

Aaron

Yeah. One thing I would add about the role itself is that I think one reason it has been successful despite Jane Street having been doing a lot of education before this role existed is having someone where that is a sort of first-class focus of their job, is sort of a step change in how they’re going to be able to engage and the sustained attention on education. But yeah, somehow there’s just more and more stuff we need to teach. And not just to new people who need to learn OCaml, but it turns out we’re changing the OCaml language. And everyone at the company, no matter how long they’ve been here, needs to learn about it. There are all these new AI tools. We want to teach everyone how to engage with these effectively. And it’s sort of the case that the more we have put into this and expanded and refined it, the more places we go like, oh, we should make that better too. And so the appetite for this kind of focused educational expertise seems to have only grown.

Ron

Right. And so on the OCaml side, we’ve posted this OxCaml educator spot for people who have a strong PL background and are also excited about teaching, and we’ve actually had some good luck on hiring there. We’re still looking for great people to do a kind of machine learning educator role where focusing on we’re doing all this exciting new generation and training of exciting new ML models and people who understand that space and can teach more people to do it is a very valuable thing. And just the general version of this role that hits all sorts of things. You mentioned having to teach everyone anew how to use the exciting new corners of OCaml as they come into existence is one thing. Also, we now do a bunch of Python, and we kind of want everyone to know something about Python. In fact, one of the things I worry about is people getting captured by their tools.

One of the things that’s nice about having a small number of programming languages is you can go anywhere into any code base and do stuff. And once you have two or three or four programming languages, there’s a kind of natural striation where people get good at the one thing. And I’d really love to be more in a world where there are no OCaml programmers and Python programmers. They’re just software engineers and they know how to use many tools and can figure out what’s the right tool to use given the particular circumstance. But for some of that, we actually have to teach people. There are people who’ve been here for a long time who’ve like, I mean, everyone knows some Python, but there are lots of people who aren’t any good at it and there’s real things to learn to be good at it. And so that’s another corner where I think there’s more teaching.

Plus we also invent our own programming languages in spots and teaching people about that and a million other things. So yeah, I do think that there’s just more surface area to cover, more technology for people to understand, and therefore more opportunities to create real value by teaching people stuff.

Aaron

Yeah. And to say another thing about this general dev educator role. I think a model that I feel like we’re heading toward and that mimics some of the ways that say designers have been incorporated into different teams and spots throughout the firm, I think that may be where we’re headed for educators as well. So I mentioned AI tools. I think we would love to have an educator sitting on our AI assistance team to focus on the education work that they’re doing. And I think more and more places will want an educator embedded there for all the teaching materials that need to come out of that part of the world.

Ron

Right. And I actually feel like there’s two benefits there. One is the thing you’re saying of there are education needs that are focused on what’s happening in a particular area. And then also I think you want the people who are doing the education work to have a variety of different technical experiences so that this is sort of more of a body of experience. If you poke them all in the same corner of the firm, they’ll have all the same context and that’s just less of a robust foundation.

Aaron

Yeah.

Ron

Let’s pivot for a second to talk about the technical work that you do. So where in Jane Street’s technical organization are you and what kind of engineering problems do you think about there?

The Editors team, and what a “feature” is (00:31:18)

Aaron

Yeah, so I’m on the Editors team within the Tools and Compilers group. So our team maintains the text editors that folks at Jane Street use to write code, VS Code, Emacs, Neovim, as well as some related services like the code indexing service that we have, the OCaml language server. And I have sort of worked across various things in the dev tools space. Started out working on an extension to our feature management software, Iron, to make it much easier to split up work into different pieces of code review and organize it in a way that was sort of awkward and manual before, but giving people a nice tool to just be like, take this and split it over here and arrange it in this nice chain of relationships that’ll all get reviewed separately.

Ron

It’s maybe worth saying when we say the word feature at Jane Street, we mean something that’s roughly analogous to sometimes analogous to commit and sometimes analogous to PR in the other world, in the outside kind of world’s dev tooling. I think the thing that really strikes me about all of the tooling we’ve built in this kind of code review and feature management area is just like it’s actually made a very different set of choices. We’ve been off on a 25-year alternative universe adventure of building development tools that just look pretty different. And if you look at Iron, which is the tool that’s responsible for code review and releasing features and all of that, it’s way better than everything in the outside world and also way worse depending on exactly which thing you want to look at. And the thing you’re talking about, this having nice ways of carving up and breaking down a larger collect change into a sequence of smaller changes is a good microcosm of, first of all, it has really good support for what we sometimes call stacked PRs in the outside world.

And that’s built in and a really nice foundational piece. And also the process of really editing difs and moving things back and forth was awful, slow and painful and the tooling wasn’t very good. And you would’ve been much happier using, I don’t know, the Emacs Magit extensions or whatever in Git than using the stuff that we’ve had. It’s just interesting that there are lots of different technical paths you can take towards the same big picture goal and end up with really different sets of choices and really different functionality as a result.

Aaron

Yeah. And there’s a whole engineering culture that flows from that. How you do code review, how you think about giving people feedback on their work is entirely sort of shaped by the underlying technical system.

The editor as home base (00:34:10)

Ron

Right. And actually, I think this comes to the question of why does Jane Street have an editors team? Exactly. Text editors, can’t you just use one of them? And I think it’s just the case that editors are much more central to how we work and how we review code than you might imagine. I don’t know, it’s worth a little bit saying how are editors here different than maybe people have seen in the outside world in terms of the role that they play in our software development process?

Aaron

Yeah, so there are a few different things. One, at a fundamental level, Jane Street has approached the text editor as the software engineer’s home base. All the interactions with Jane Street’s code review and version control and all these other systems live within the editor. You do everything there. It’s all easily keyboard-navigable, and you’re not sort of jumping between different applications or pages. And so that means that because Jane Street does a lot of weird different stuff, like not using Git, but using Iron, we actually need a team to build all this integration into the editors so that people can actually live there and not do some things in the command line and some in the browser and then write code in the editor and sort of jumping between these different surfaces.

Ron

And despite the fact that we’ve built a lot of, I think, quite cool web UIs over the last two years and tooling for building web UIs, Jane Street is just a much less web-focused place certainly than the big tech firms, which I think have all ended up leaning quite heavily in terms of taking all of these interactions and sticking them into a web browser. And for us, yeah, it’s just like you’re in the editor, you make the changes, and the editor becomes both the place in which you interact with all these tools and also the central location for understanding code and for communicating about code. I think another interesting early design choice, which has ups and downs, but I think has certainly cemented the importance of editors, is that the way in which you give feedback on a change is in the editor. There’s specialized comments that you write with a particular format to mark a piece of feedback rather than going off into some separate web UI and clicking a button and leaving a note on a kind of PR that somebody else owns.

And part of it is this formal feedback that comes through the editor, but also because you’re in the editor, some of what you do when working on a feature with someone is you just collaborate on the feature. Instead of leaving a CR, which is the name for these comments that we put in, you might just fix the thing, you might just make a change, which I think is a great. I think it’s a great part of how development works here. It’s also weird and unsettling, I think for people who are used to the other thing. It’s like someone else is reaching in and modifying my PR. It feels unsanitary. But it’s just for good or ill, or maybe I just want to say for good and ill, it’s just a different modality, a different way of interacting with and thinking about code. And the editors just have a really central piece in all of that.

Aaron

Yeah. And as you’re saying, this really biases us to keep the editor as the home base, because if you’re interacting with these systems in a surface where you can’t just as easily jump in and edit the code, you then are sort of losing that property of this code review system and property of a code review culture that I think we see as pretty positive.

Giant diffs and other problems of scale (00:37:47)

Ron

What is challenging about this at a technical level? What are the things that make it hard to make the editors do the things that we want them to do?

Aaron

So one aspect is that the editors that we’re using are things that exist in the outside world, VS Code Emacs. And for example, one problem that we’re at work solving in VS Code right now is that in JS work, we sometimes have incredibly large diffs. So a giant configuration file or just a massive OCaml file. And it turns out that the diff algorithm in VS Code has some serious performance problems once you have diffs of a certain size. And there’s furthermore an issue of if you try and open a large diff and it hangs and you close it, it’s now locked up the queue for viewing diffs at all.

Ron

Oh no.

Aaron

And you can’t look at any others. And so there are just things like this where either there’s some functionality that we want internally that isn’t accommodated by this external tool that we’re building in.

There are also challenges where we have implemented some integration with the Iron code review system that messes with some assumptions that say VS Code makes about what a buffer is going to be, what the URI for that buffer is, and then say makes the built-in editor search misbehave in a strange way, and we have to track that down and fix it. I think that there are also challenging things about the shared infrastructure between the editors, like the OCaml language server and the underlying system Merlin that provides type lookups and jump to definition for OCaml. And there, there’s lots of interesting performance problems and questions about what are the optimal data structures. And if we want to build an index of the entire monorepo, oh, someone has added extra information to do better – to the files that the compiler produces to do better lookups of modules using this index.

Well, now this index uses many gigabytes more of RAM and we can’t actually build the index on our process anymore. As I’m saying this, I realize a lot of the problems we’re trying to solve relate to the scale of certain things that we’re doing at Jane Street and making these general purpose tools perform really well at these idiosyncratic use cases.

Python notebooks (00:40:41)

Ron

Right. And I think one thing that comes out of what you’re describing here is that the editor has this kind of central role where it’s like the surface through which many other things come at the users. And so maybe the problems you run into are problems about the type lookup system or this indexing system, things that are not literally part of the editor, but the editor is the place where you see the issues. And so you’re driven to fix things sometimes in the code review system and sometimes often the OCaml language stuff. Increasingly the Python world is another place where we’re spending a bunch of time. Actually, maybe it’s worth saying what has changed as we’ve needed to do more support for Python? I’m curious what kind of pressures that puts on the editors because in some sense it should all be easy. The whole world uses Python.

Now we can just do what everybody else does.

Aaron

I think the main surface for folks in trading and research who are interacting with Python is the Python notebook. And this has been I think the area that’s been very challenging. There are different notebook solutions. They all have different significant flaws or limitations. And I think it’s been one of the pressures is even understanding what the different use cases are and why people use one over the other and a sort of unfortunate trend of spinning up a new one that we hoped would replace the old one, except it’s missing some functionality that the old thing has, and so now we’re maintaining multiple of them. And one other pressure is that for people who are writing a lot of code in Python notebooks, that code is not always easily shareable. And for example, you might want to take it and put it in a Python library that lives in the repo that can be imported and shared between multiple notebooks.

And for that, you’d really want people in a surface like VS Code where it’s sort of just as accessible to edit Python files as it is notebooks. But for at least the kind of Python notebooks work that people do here, VS Code has been very challenging and has had a lot of performance and usability problems to the point where we are now starting to work on just building our own solution to this, which is I would say a theme of a lot of the tools work at Jane Street.

Build it ourselves, or push it upstream? (00:43:30)

Ron

Right. And I do worry a lot that we overlean in the direction of building our own things. At the same time, there’s a lot of power in having control and having everything built in an ecosystem that you understand from top to bottom. So I actually feel like you see work happening, especially in the editor space, that’s shearing in both directions. There’s some places where in the notebook work, we’re like, oh yeah, now we’re going to, instead of these multiple things on the outside world, none of which exactly work in the way that we want, none of them have quite the design goals that we want. Some of the ones that look closer aren’t maintained anymore, whatever. There’s a whole bunch of problems with the kind of notebook stuff in the outside world where we’re going to write a nice little OCaml thing, use our own web UI stuff, and it’s just going to work smoothly and we have a prototype that’s pretty good.

And I feel pretty optimistic about this being both not that hard to land and also giving us a huge amount of control to be able to adapt the notebook to the very specific things that we want to do internally and have a much better experience. At the same time, I feel like a lot of the other work that you guys have been doing is about de-weirdifying various aspects of our setup. We have 25 years of weirdo special case Emacs stuff that we have built up over time. And we’ve just been spending a lot of time funding work on Emacs, figuring out how we can take the weird custom thing that we’ve done internally and replace it with the standard open source thing. So I feel like actually both of these directions are good. You just kind of have to pick your spot.

Aaron

If you can find a way to make the upstream thing do what you want, that’s just so much nicer. Other people improve it for you.

Ron

The code that you don’t have to write is the best code of all.

Aaron

Not only are you not writing, but lots of other people are using it and reporting problems and those are getting fixed. We don’t have to find all the problems ourselves by running into them.

Ron

And it turns out the upstream Emacs people really understand Emacs well and they know how the system works and they have the design aesthetic down. And in many ways you just get better stuff than you would get from our custom creations.

Aaron

Yeah. All the different pieces work together better that some external package will play much more nicely with our internal extension if our internal extension does it the sort of native Emacs way as opposed to some totally different approach that then breaks when it comes into contact with the outside worlds.

Ron

Right. Although a thing that pushes in the opposite direction is a bunch of the extensions that we’ve done, we’ve written in OCaml, which is super not the standard way. In fact, we do this across multiple editors for Neovim and Emacs and VS Code, we have OCaml as at least one of the major implementation languages. You also have to use the kind of more native language to do some of the pieces. So how does that play into the technical tension?

Aaron

Yeah, it definitely complicates the story. And in Emacs in particular, we have some ideas about we haven’t put the boundary between Elisp, the native Emacs implementation language and eCaml, the flavor of OCaml that we use to write Emacs things. The boundary there, particularly for asynchronous things, is pretty awkward. And so I think a long-term goal would be re-architecting that to have fewer potential for weird deadlocks when these two sides are waiting on each other or things like this.

Ron

Is part of the issue here is that just Emacs is super ancient and does not have a great concurrency story at all?

Aaron

Yes. So I think to some extent that’s true. I think upstream Emacs is actually working on this. I don’t think they, and in particular, there were some changes to the garbage collector that I think may play more nicely with asynchronous things, but unfortunately that wasn’t ready for the current cut of Emacs that we’re working on upgrading to. Emacs and Vim are both sort of tools that have been around for a while and have some limitations in terms of the kinds of interactions that they were intended to support.

What agentic AI changes for editors (00:48:05)

Ron

And even beyond that, just marrying the concurrency primitives of two different languages is really hard. We actually have run into very similar challenges around Python and OCaml where I think we have now maybe started to work out a story that makes more sense, but it’s a tricky thing even when you have pretty good concurrency stories on both sides. So to pivot a little bit, I think a thing that has affected everything about developer tools is the emergence of actual good agentic AI for programming. And I’m curious how that has affected the way you think about work in the editor and what you prioritize and what are the things that you want to build.

Aaron

Yeah, there are a few different pieces to this. One, I think it hasn’t really affected our belief that the editor can still serve as an excellent home base for the software engineer and that we want to make the new workflows that are emerging just as natural and accessible via the editor as the ways people were working before. And then there are also some things that seem like they aren’t going to necessarily change about what people are doing in the editors. This is looking at diffs, doing code review, working to understand and read code, whereas how people write code may be changing and is changing significantly.

Ron

There are certainly plenty of engineers all over the place, but also at Jane Street who just don’t type code very much anymore. And most of the typing of, they type more in the way of prompts, but they’re still reading code and they’re still looking at the types of things and looking at the definitions of things. So they’re doing lots of things you do in the editor, but maybe less editing.

Aaron

Yes. And this I think has created quite a lot of uncertainty for people working on editors in terms of what is it that is useful to build now? What is it that is going to be useful a year from now? Will anything? And one thing that I’ve gotten quite interested in recently, particularly because the editors team has been investing in improving the sort of telemetry and observability of what people are doing in editors and how they’re interacting with the OCaml language server, is trying to use this telemetry to understand how is it that people’s work is actually changing? I looked the other day, I had a hypothesis that, well, if I look at the number of times per user that people jumped to definition in OCaml, if the LLMs are producing good code and people aren’t writing and having to look up APIs as much, they’ll be jumping to definition less on a per user basis.

But then in the telemetry, it seemed, if anything, they’re doing it slightly more, which was quite surprising, but also maybe relates to people are actually needing to read and understand more code than they were before. And so I’m excited for this kind of data to guide, to help us understand what is it that people are actually doing in editors because it’s kind of been quite amorphous as things have been changing very quickly and using that to help us decide what is important to invest in and what are the sort of workflows to make better because questions like, should we try hard to make jump to definition faster? That depends if people are using it more or less.

Ron

Yeah, no, it makes a ton of sense. And the thing you’re proposing of using the data to drive seems more right than my instinct, which is theorizing the air about it. But on the theorizing the air part, it does feel like understanding is at the moment, the thing that’s most important. The cost of creating a plausible looking PR. I think I’m stealing something from Charlie Marsh, I think. The cost of making a plausible PR has dropped nearly to zero and the cost of verifying it is about the same or maybe worse. I think notably there’s something kind of uncanny about a lot of the features generated by LLMs in that they’re very smooth. They look very nice. A lot of tells of things that you might think of as things that are indications of a problem just aren’t there, but they’re often still super broken.

And so the process of reading through and understanding is at a minimum different and in some ways certainly harder. And so the idea that we need to spend more energy and engineering work making it easy for people to understand what is happening seems really important. And I think some of that understanding should come from further leveraging of LLMs. I think LLMs are a really powerful explanatory tool. But I think it’s easy to underrate the value of traditional hard ways of extracting data from code and surfacing it. I think things like go to definition and inferred types and test results and things like that are all really exciting.

Claude Code’s learning mode and mental models (00:53:27)

Aaron

Yeah. I’ve also been thinking more about people’s actual sort of workflows when getting LLMs to write code and how that relates to understanding it at the end result. So one thing that I have played around with in some of my education work is Claude Code’s learning output style, which changes the system prompt to tell it stop sometimes and leave a to-do human sort of comment in the code and ask the human to go fill in that piece. And in using this myself, I found that having it stop partway through, and then I had to read what it had written so far and understand it enough to fill in this piece. I felt like I came out at the end, both have had a sort of more pleasant experience of producing this code, but also a more pleasant experience and I think effective in actually understanding what it had produced.

And I’m curious to explore not just on the output end, adding more explanatory power and tools to the code that’s already produced, but thinking about how do people interact with these models as they’re producing the code to most effectively end up with a solid understanding of it at the end and not just be guided by, well, it feels more efficient to have it write everything and then read it at the end as I’m become a little skeptical that that is actually sort of the optimal way to approach it.

Ron

It’s interesting. It reminds me a little bit of the thing we were talking about, about the utility of taking notes while writing with a pencil. The idea that you are involved in the production function will cause you to better understand the thing at the end. And there’s maybe a kind of infelicity in the current way the tools are set up in that they largely. Everything is kind of tilted in the direction of the model generates the code and then you decide whether or not to accept it and maybe that’s wrong. I thought, that’s a fascinating idea.

Aaron

Yeah. I was giving a talk to a group of interns about code review recently and had a slide. Code review is hard. Writing code, you’re sort of incrementally building your understanding. And when you’re catching mistakes, you’re sort of correcting your mental model. And when you’re reviewing someone else’s code, they’re just presenting you with the finished result and you have to awkwardly build your own mental model of it. And it immediately struck me, wait, this is exactly what working with LLMs is like in the way that it makes it harder to have a mental model of the code you are authoring. That sort of started me more firmly down this path of, yeah, maybe we should be interacting with these things differently, at least some of the time.

Ron

That’s really interesting. There’s just a lot of space for innovation here and new methods of approaching. Just like LLMs suddenly make the things that you can do. It’s a very different cost structure and there’s a very different opportunity set. And so the thing you’re talking about with the Claude Code learning mode thing is something where you can integrate a kind of tutor-like behavior into the automation and that’s just dramatically cheaper of a thing to have done than before you have to go off and go and talk to another person to get that experience to happen. I also feel like in the editor space in general, there’s a lot of room for custom integrated visualizations and ways of understanding the data and that the models are really good at ginning these things up. And the idea that we should be able to ship along with a feature a custom visualization that you made for understanding that diff that just kind of shows up in the middle of review.That’s a thing that we could plausibly do now that would’ve made no sense three years ago.

Aaron

Yeah. And one thing that I’m excited to explore that connects to your point about these things being very good at explaining is that you could also bake into a feature some sort of pre-computed segmentation of the code into logical chunks and then a pre-baked explanation of each of those chunks that you can just call up with a single keystroke as you’re reviewing the code. I mean, as you’re reviewing it, you can highlight and ask a model to explain the piece, but that’s some additional friction. The response doesn’t necessarily come back that quickly. And if for any given piece you could call this up, I think that that might also be a sort of powerful tool for helping people more easily understand the increasing volume of code they’re being asked to review.

Teaching when the LLM can do the exercise (00:58:20)

Ron

Right. And I think this is just the modern condition of a software engineer, and by modern I mean for the last eight months or something, is that there’s just a lot more review to do as a proportion of your time. You’re just spending much more time reading and understanding. And so suddenly the value of making that better just shoots up by a lot. I mean, it was already a lot of work, so it was already somewhat valuable to do it, but it’s now much more clearly the rate limiting step in all of this work. So we’re talking a bunch about how AI affects the experience of developing code and understanding code. And this resonates, I think, a lot with the question of how does AI affect the process of teaching people stuff? So I’m kind of curious what that has felt like as you’ve been helping to run education programs here amidst the middle of this AI revolution.

Aaron

Yeah. So this is, I would say, quite a thorny problem because you’re faced with this tension of. One of the things that I help work on, and maybe we can talk more about this, but this workshop for early career engineers at Jane Street focused on techniques for software testing. And we call these workshops teach-ins. So this testing teach-in teaches people how to interact with various libraries for property-based testing or for deterministic control of time in tests. And when you’re teaching someone how to use an unfamiliar library, well, an LLM could just write that for them. So what is it about this library that someone actually needs to know in practice and what might you consider a boilerplate that you don’t need?

Ron

The initial scaffolding can be made really easy, right? The models can do a decent job of that, but what’s the understanding that you have to get to people?

Aaron

Yes. And so I think so far we’ve come down on the side of people having to think through these sorts of problems and implement them themselves at least once is actually pretty important for coming out of it with a sufficient understanding. And even if what they are doing subsequent to that is reading applications of this technique initially authored by an LLM, building that understanding by having done it themselves once is sort of a critical input to then being able to evaluate those things. That is one piece of it. And I think we are actively thinking about where to draw the line in terms of, there’s some stuff that we were probably asking people to do in the past because it was sort of stuff that they had to do even though it was somewhat toilsome that yeah, we should probably not have people try and learn how to do those pieces.

But there’s also, I think the challenging question of how do we teach people and even what do we teach people about using these tools effectively, about how to integrate them with their workflow, about how to understand the ways in which they can be effective or lead you astray? And I think this is still in very early days. I don’t think we have sort of settled answers to this and the tools are changing very quickly.

Ron

But you can already see problems. There are definitely people who run into problems where they just kind of vibe it up too much and use the LLMs to generate stuff and do not think. And that’s a real problem. You still need to think. I actually worry more generally about what it means for the overall education system. I think a kind of critical component of learning is struggling with the material. And the world where you always have a genie at your beck and call that you can just summon up to solve the problem is itself a real problem. And the overlap of problems that are easy enough, that they are suitable for someone who is learning to do and hard enough that an LLM can’t just one shot it is close to the empty set now. So it’s like a really challenging problem. I mean, it really does make me just think back to the doing stuff on paper without a computer.

Sometimes an important thing about making the educational process work is withholding certain affordances. And I worry that we have a whole generation of people who are learning in this new environment where we don’t know what to do. We don’t know what affordances to provide and which ones to withhold. And I worry that we’re not building a good educational system for people in this world.

Aaron

Yeah, I think you and pretty much everyone else involved in education is deeply worried about this. To your first point about how do you get something that’s small enough to be tractable for students and large enough to not be trivial with an LLM? We have run into this exact problem multiple times with new curriculum we’ve built and we’ve sort of repeatedly ended up, the thing we’re asking people to do is too large for the time we have. We are asking them to do something where they do have to put a lot of work into the LLM code to make it good and actually do what they want. But that just takes too much time to do well given the amount of time we have on the education piece. And I think it’s going to take more experimentation to figure out where that lands. But yes, I think that figuring out where to slot using these tools into sort of traditional education of it feels inevitable that we’re going to end up with some sort of structure where you do this piece or this set of activities without the aid of LMs and then you do some with.

And I think figuring out how to make that pairing and what you’re trying to get out of each of them is something that a lot of people are going to be experimenting and figuring out.

Ron

Yeah. And I think it’ll be interesting over the next few years as we’re doing this increasingly with cohorts of people who have not spent that much time typing code themselves. They’ll just be like, there’s a core learning mechanism that maybe they haven’t had enough experience with. And yeah, I wonder if that will change what we have to teach people in the end.

Aaron

Yeah, there’s been a lot of conversation about things along the lines of that Claude Code learning mode where what LLM tools, if any, should we provide to folks going through OCaml bootcamp, learning OCaml for the first time? And if they’re just having LLMs write all the code, well, it doesn’t seem like they’re going to learn anything. But it also feels a bit odd if people going to leave this and then LLMs are going to be a core part of their day-to-day work. And so not having those be at all a part of this also feels like maybe not a great match for what we want to teach them. And so can we come up with prompts, something, a mode where the LLM will just sort of ask you questions to check your understanding or we’ll leave you large pieces to do yourself. We’ll sort of set up a kind of skeleton, but actually ask you to do most of the implementation.

And I think there’s going to be a lot of experimentation with that as well.

Ron

And maybe the thing that’s been talked about a lot in educational circles of setting up the LLMs to be effective tutors so that there’s a part you’re supposed to write yourself and maybe you get stuck on something and maybe you can get the LLM to give you a hint, but not to go off and write the code for you so that you do have the opportunity to actually do the hard thinking yourself.

Aaron

Yeah, this is actually a category of thing that I did research on in grad school of an automated tutor that could walk you through an arbitrary example problem where you had some algorithm that you wanted to teach people to follow to solve this kind of problem. And with the idea that sure, you had the one hour where you could go talk to a TA and get one-on-one help and that was the most effective thing you could do. But if you were at other times studying for an exam, it might be really useful to have this not as effective, but still helpful sort of tutor that is infinitely patient and can walk you through any example. Yeah, I think that that is likely going to become a large part of at least higher education.

Bootcamps, teach-ins, and a course catalog (01:07:26)

Ron

Makes sense. So we just talked about this testing teach-in, which is part of the kind of formal educational program that we have for new engineers. Maybe it’s worth walking through what do we do to onboard people and teach them about our world?

Aaron

Right. So a new software engineer joins and the first thing that they’re likely to do is OCaml Bootcamp, a sort of cumulative series of exercises, building up a little client server app in OCaml and learning the different core libraries and getting their code reviewed by a mentor and starting to understand some of Jane Street’s style. And from there we have something called a production bootcamp where you go through a series of exercises to learn how applications are actually deployed and monitored and configured at Jane Street. And that forms say the first two weeks someone is on the job. And then -

Ron

And to me, it’s worth saying that’s really been refined a lot over the years. I think the OCaml bootcamp itself, early versions of that are something like 20 years old. But there’s been a lot of work. In fact, you’ve been involved in a bunch of the recent stuff to refine and hone and make that better. And the fact that nowadays someone can arrive and get through that in about a week and then a little more for the production bootcamp is already a really good improvement. I think we’ve been able to make that process more efficient. And I don’t know, people who come into Jane Street, there are good and bad things. People aren’t, for example, always super enthused about the overall state of all of our documentation, which has its ups and downs. But I feel like I’ve gotten, especially in the last few years, really positive reviews from people of these early bootcamps.

They really, I think, do a really good job of setting people up with the basics that they need to do and doing it in a really efficient way. So I think that’s been a nice piece of progress.

Aaron

Yeah. One thing that we’ve worked hard on, which I think is important when teaching this kind of material is introducing new concepts in context where it’s motivated by some specific application, some feature you want to add to this system you’re working on. And I always frame this as the mistake people make is they introduce a tool and then go in search of a problem that it solves rather than like, “Hey, we’re working on this thing. Now we want to do this. We don’t have a good way to do this, but wait, here’s this thing. Why don’t we use this?” And yeah, I think that makes it smoother. That helps people have hooks that they can hang their understanding on. And so that’s been a lot of the work to rework these intro materials.

Ron

Cool.

Aaron

From there, in someone’s first year or so at the firm, they’ll go through this series of teach-ins. So I mentioned the testing one. There’s one on writing performant OCaml code and the tools to analyze the performance of OCaml code. There’s one on systems debugging, all the different tools we have for understanding what might be going wrong with some system. There’s one on advanced functional programming, a bunch of the advanced features of OCaml and cool ways to apply them. And then there are also some more specialized concepts. So a teach-in on Bonsai, the library for writing web pages and web UIs via OCaml. There’s one on market data, how the systems for market data work at Jane Street, which is very important to some people, but other desks not so much. And I think as we add more, we’re going to be even more in the mode of it’s sort of a course catalog.

And you and your manager or folks on your team sit down when you start and think about which of these are relevant to you to do when and almost devise a course of study for your first year at the firm. And we also try and create these materials in a way that’s amenable to self-study. And you could also go and work through them on your own at the point where you actually need to learn and apply the material.

Ron

One of the interesting things about the teach-ins and about how those have evolved over time, they were initially built as a kind of reflection of a trading side education program also called Teach-Ins, where that was mostly organized around basically going around from trading desk to trading desk, learning the kind of core concepts that were important for that desk. And my mental model of it early on was there’s some specific technology you want to teach people about, but the more important thing is the kind of conceptual layer of Bonsai is a web framework, but it’s also an introduction to incremental computing. And market data is a particular domain you need, but it’s also thinking about handling complex streaming data and building modular and efficient protocols for consuming those in a good way. There’s a bunch of architectural lessons from each one of these. And I feel like that initial vision sounds good and has been less successful than I expect it to be.

And we’ve over time gone in a direction that’s been more about giving people the practical tools that let them go do things. I’m kind of curious, how do you think about that? How does the conceptual piece fit in? When and where should we be focusing on that? And to what degree should we just be giving people what feel like the most practically useful tools?

Aaron

Yeah, this is a very interesting difference in the Jane Street context compared to say college courses, which in a college degree, maybe I’m teaching a databases course and there’s some domain knowledge about databases I want to impart. But as part of their university education, I also want to teach them to be capable problem solvers and how to work through ambiguity or try out different solutions, a lot of these general skills. And one of the great things about working at Jane Street is that my colleagues are very smart and capable and I don’t really need to use their time to teach them problem solving skills.

Ron

They’re already really good problem solvers, it turns out.

Aaron

There’s some selection effect there. And so this I think has biased the education away from what you might see in university of we’re going to have you spend a bunch of time grappling with this tricky problem and figuring out a solution. Because actually when people are at their desk in their day-to-day job, that tends to be what they are doing. And it therefore can feel pretty, I guess, unsatisfying to do that on a kind of canned educational thing as opposed to something that’s more directed, here are useful tools that we’re going to have you work through and understand, which you would then go apply to the truly tricky problems. So I think that you are right that there are some core conceptual things like the incremental computing, things like this, which we do want people to be exposed to and think about, but we tend not to do that by asking them to take on some particularly sort of thinky piece of it to spend several hours just having to come up with a kind of novel solution to it.

Ron

What you’re saying here is even if a big part of your goal is to convey the conceptual lessons, it will work better if the way you do that is by teaching them concrete, useful tools and focusing on those. And some lessons will sneak in that way, but not because you shaped it explicitly about that, but because understanding the tool and being able to use it required the understanding of those lessons.

Aaron

Yeah, I think that’s right. And also teaching this conceptual knowledge just in a different way than you would in say university setting and doing it often more in ways that look like a walkthrough of let me take you through this kind of problem and how you would solve it. And you’ll do some pieces, but how you will build the more complete understanding is when you actually go and apply this to the problems you’re solving on the desk.

What are the learning goals? (01:16:15)

Ron

So this is a little bit about some ways in which teaching at Jane Street is different from teaching the outside world. I feel like a thing I learned from you early on is about some of the ways in which lessons from the outside world apply here. So we talked a bunch about the testing teach-in. This was actually the first real project that you did here where the way I remember it is we were trying to organize this testing teach-in before you started and it was going badly. A little bit because it was this problem of it was hard to quite get enough of anyone’s time to. People were doing good work and focusing on various parts of it, but there wasn’t a real effort to bring the whole thing together and make the whole thing good and land it. And I remember you asking me early on, because the shape of it was a thing I had come up with.

And so you asked me, so what are the learning goals? And I was like, fascinating. What are learning goals? It sounds like those are English words that kind of make sense together, but clearly you mean it in a more specific way. And I found that to be in the end, a very useful and insightful way of thinking about teaching. Maybe can you say a few more words about what are learning goals and how should that affect how you design curriculum and build courses?

Aaron

Yeah. So when starting off on any larger project, it’s good to write down what you were trying to accomplish. And in this case, learning goals are a specific way of writing down what you want to accomplish in terms of what is changing about the students as a result of the thing that you’re doing. And specifically, I find it very useful to write them in the form after completing this teach-in, say, students will be able to and then write down a bunch of active verb phrases. So for example, understand property-based testing is not a great learning goal because it’s very vague. It’s not a specific thing that they’re going to be able to do.

Ron

And how would you measure it?

Aaron

And hard to know how you would measure it, but students will be able to write a property-based test, like a sort of nested record type input. So some sort of concrete scoping, what is the complexity and the sort of thing they’re going to be able to do. One that really forces you to structure your thinking about what am I actually trying to accomplish here? And importantly, what am I not trying to accomplish? What is not a good use of students’ time? Because it’s not actually a goal to have them do this thing. But then in terms of the overall design of the educational program, once you have written down pretty specifically what you want people to do or what you want people to be able to do afterward, well, it follows pretty directly what you should have them do in the thing that you’re running because they should practice the things that they’re going to be able to do when they’re at the end. And okay, now I know the sort of exercises or activities I’m going to have them do. Okay, what is the background that they need to know for this? So what lectures or other materials or what parts of code would we want to provide them for which writing those is not relevant to these goals? And so the whole process of developing some new piece of curriculum flows nicely out of having been disciplined in writing down these goals.

And I feel like to some extent, my education work at Jane Street has been going into different rooms with different people and saying the word learning goals. We should do this.

Ron

And the other thing that I found interesting about the learning goal idea is it tells you something about how you should think about evaluation. And here, we mostly are not so interested in our teaching stuff to evaluate the students, but you really want to evaluate the program of did the thing you do actually teach people anything useful? And so the idea that you could see whether or not at the end that people are able to do the thing that you wanted to do is a way of evaluating yourself and your own work as an educator. And I guess the other thing that I think is striking about is I think that learning goals has now showed up for me as in some sense two different things. One is there’s this somewhat more specific way of writing them out and thinking them through and using them to structure the program.

And then there’s also just the frame of mind of how are the students going to be different at the end? What are you even doing here? Why did you show up to work today? What are you trying to achieve? And I found actually that framing to be very useful in talking to potential developer educators. Talking to people who you want to hire to do education work is I think some people have really thoughtful answers of what they were trying to achieve in the classes that they taught. And sometimes it’s like, oh, I just like this material, so I taught a course about it. And I don’t want to complain too much. I think a lot of great classes are just like someone is excited about material and they teach it. And an enormous amount of good education is not about process, but is about charisma and ideas and being able to just carry the performance through and really caring about the end results.

But I do feel like the structure helps and especially helps people who haven’t done it before. People who are trying to figure out how to teach classes for the first time. I think it’s very helpful way of organizing your work and your thoughts.

Aaron

And it forces you to focus on things that are not just in the student’s head to some extent of what am I trying to accomplish here that isn’t just sort of like I tell the student a thing and now they know it. But yeah, how are they changed in some sort of observable way?

Ron

So are there other aspects of the outside learning curricula and approach? The whole world of research on education is complicated and messy and also includes lots of conclusions that probably aren’t true. And also lots of conclusions that we’ve known for a while that somehow don’t get applied. There’s all this crazy stuff about ways of teaching reading where we had pretty good research telling us what to do and then a lot of not doing it for all sorts of messy reasons. And yet I think the learning goals is a nice example of there being real insight that comes out of the whole intellectual stew of the educational world. Are there other pieces that you think of as really important and valuable that you try and leverage here?

Aaron

Yeah. One that I’ve tried very deliberately to apply to the OCaml bootcamp is what is sometimes called cognitive load or you might call it working memory. But just when you’re trying to learn something new, there’s actually a limit to how many new things you can handle at the same time before it becomes harder to learn any of them. And so in an older version of the OCaml bootcamp, right away you are learning OCaml, you were learning an editor that perhaps you hadn’t used before, and you were learning a version control and code review system that you also hadn’t used before, and you were doing all of these from the jump. And one thing that I did was sort of split these apart. So you start off just learning OCaml syntax in utop, the sort of command line REPL. Now we have nice in-browser text editor boxes where you can run little snippets of OCaml.

And then you go to a piece where you get introduced to your editor and you do some activities just with the version control system. And only after you’ve done each of these sort of introductory pieces do you get to the part where you’re bringing them all together. So applying this cognitive load idea in that way. And I think the other thing that comes to mind, which I mentioned before, is this idea of active learning, of trying to make the educational activities not just someone telling you things, but as interactive as possible. And I think that some of the new material that we have developed for the internship this summer is really applying this in some interesting ways of lots of small group discussions or interactive activities around design or practicing talking to stakeholders. Yeah, applying this idea of people actually interacting with other students as an important part of the educational process.

Ron

Right. Again, leading into, like, this isn’t just about conveying information. There’s a big psychological element to people learning and retaining.

Aaron

That’s right.

Rebuilding the intern curriculum for AI (01:25:10)

Ron

So maybe before actually talking for a second about this point about the internship, so you’ve been involved in what has been a pretty serious rework of the educational intro that people get to the firm. Can you say a little bit more about what motivated that and what we’ve tried to do in that space?

Aaron

Yeah. So the precipitating motivation was, wow, AI tools are really good now. And maybe the four-week project we used to have interns start off by doing, that’s one prompt and you’re done. And so that’s not a useful way to evaluate an intern or productive work to ask an intern to do. And so, okay, we’re going to change the nature of the projects that interns are going to do to be more open-ended, to maybe involve more elements of design or gathering requirements, things like this. Well,

Ron

Basically making them harder and more realistic.

Aaron

Yes. And more going beyond can you produce a reasonable amount of good code? Just producing the code, that used to be maybe among the most important signals and it’s less important now, I would say. And so aligning the projects with these are the things we want to see people do successfully. And well, to set them up for success, we should try and teach them some more about how to approach this, how we think about this at Jane Street. And in addition, not just dump them into the deep end of you have access to LLMs, good luck, but actually try and train them some on what to keep an eye out for, what are effective ways of working with these tools, what tools even exist at Jane Street, and how do people use them? So the new materials have been a combination of giving people, and this is the sort of balancing act that I was talking about earlier, tasks that are sufficiently complicated or large that the LLM doesn’t just give you a perfectly fine version on the first try, but not too large, that it’s going to take you a week to work through all the code that is produced.

And also taking this as an opportunity to focus on some new software engineering topics that maybe didn’t get a lot of formal training at the start of the internship, like testing. Writing a very good test suite for this database-backed application is part of the new curriculum. And that’s

Ron

And testing has always been important, but in many ways more important now because the models thrive on feedback. You really need to give them ways of figuring out whether or not the thing they’re doing is right. They’re not just going to, when you get beyond relatively simple things, they’re just not going to one-shot good answers.

Aaron

For sure. Yeah. It’s been important, but there’s always a balance of like, well, we could spend half of the internship teaching people things, but we actually want them to spend most of their time working on a project. So in this new world, we’re like, okay, we’re going to definitely invest more in education, but we still at some point need to get them starting to work on their projects.

So there’s also been new exercises around looking at design docs and judging what is a good design doc, what is looking at some bad ones and discussing the flaws and more on this sort of planning side of things. We developed a new activity where there are various full-timers role-playing different stakeholders for this app that the interns are working on, and they talk to them and try and gather what are the requirements for what I would design. And then the final piece that I would mention is that we’ve talked about code review becoming a sort of more central part of the job. And so we traditionally have not had interns review other people’s code. At least that would be pretty uncommon. But this summer, this is something that we’re going to ask interns to do and that we want to actually understand how well they perform at it.

But to that end, we actually have three additional days of training in the middle of the internship on code review where the interns will see pre-made features actually adding functionality to the same application that they were doing the testing and design work on earlier in the system. And they are leaving code review comments and then they are seeing a reference of here is the set of the answer key. Here’s the set of code review comments that someone might have left and discussing as a group sort of what they saw, are there ones that they disagree with? And then also implementing their own features and interns reviewing each other’s code to give them some reps doing this with the idea that then their mentors and their actual project will send them some features for review.

Ron

So this is a really big change, both to the educational part of the internship and really to how the internship as a whole runs. How do you think it’s gone?

Aaron

I think we’re learning a lot.

I think a lot of the challenges that have come up are a lot related to the scope of what we’re asking people to do. So if we want people to generate code with LLMs and really hammer the message in, you own this code, you need to review it carefully. You need to get it into a good state. You don’t just stamp whatever the LLM produced. Well, if we ask people to produce much more code than they have time to polish. In some sense, that doesn’t quite line up and we’re not kind of communicating the lesson that we want. So I think we’ve sort of learned a lot about what is the right scope for these different pieces? And similar with the code review, sending people too large a feature with too many problems, and if it’s hard for people to spot the issues that we actually want them paying attention to in code review.

Ron

And maybe goes to the cognitive load issues you were talking about before.

Aaron

Yes. So I think there’s been some challenges there, but I also think it has been really great to see this sort of more interactive form of education compared to say the OCaml bootcamp that the interns went through because in that they’re going at their own pace, they’re getting code reviewed by a full-timer and sort of discussing with them, but they’re not necessarily doing a lot of talking to other interns about the material. But this new curriculum, interns are in these pods of four people and it’s explicitly structured with these small group discussions frequently throughout it, both like, how are you using the LLM tools? What problems are you seeing with the sort of tests that it’s spitting out to you? And I think that those discussions, including full-time mentors, have been really productive. And I’m excited about anchoring a lot of our, at least intern education in that sort of model.

Ron

So it sounds like next time you run this, that we run this, you’ll want to reduce the scope of some of the tasks people get.

Aaron

Yes.

Ron

Are there any other changes you think you should make?

Aaron

So I think that reducing the scope, but also there are adjustments to emphasis and pacing. So one thing that is always a challenge with programs like this is that people go through them at very different paces. And how do you accommodate someone who’s racing way ahead? You want useful things for them to do, and people are going through it more slowly, what do you actually have them focus on? So some of this is about being careful with the schedule and also how we are messaging about what the expectations are, such that people are actually taking the time to learn from the things that we think are most important and not trying to rush ahead to sort of keep to some schedule. And I think at least there are different groups of interns and this whole curriculum has been evolving rapidly and significantly between each batch of interns as they start throughout the summer.

But in an early one, we had an issue of starting out with emphasis on a lot of multitasking and sort of vibe coding stuff. I think we felt afterward that that was not sort of the right place to start. It sort of set a bad kind of tone for the rest of the exercise. And there’s been some rearranging to. We’re actually going to start with the testing focus piece and making this testing suite really high quality and sort of framing it around that rather than, oh, you can have an LLM do three different things at once and it’s cool.

Ron

It is cool.

Aaron

Yeah. I was talking earlier about thinking about different ways of working with these models and having people checking in with what they’re doing more often. And I think by next summer, I think we might have a lot more considered ideas about this or people have tried different things. I think it’s very possible that we’ll want to build more sort of guidance into like, here are specific workflows to follow. Whereas I think now it was, yeah, we have some things to know about ways in which these systems can misbehave and things to keep an eye out for, but not a lot about we think this is the good way to prompt or interact with these systems. And this is another reason why these small group discussions were great because people were sharing the sort of different techniques that they were playing around with and what was working well and what sort of wasn’t serving them.

Ron

And I guess one of the challenges of this particular aspect of the world is it’s just changing fast. And so knowing what advice to give, it’s like we can give some advice now. Next year might need to be different advice.

Aaron

Not only is it changing fast, but two people using it in the same way might have it do different things. So when you give people an exercise that involves using an LLM, it just introduces much more uncertainty into what is going to happen when they do it than you would’ve had in the past.

Ron

Right. All right. Well, maybe that’s a good place to stop. Thank you so much for joining me.

Aaron

This is great. Thanks so much.

Ron

You’ll find a complete transcript of the episode along with show notes and links at signalsandthreads.com.