With all of BayLeaf’s users concentrated in one specific geographical area, following the rhythms and schedules of campus life, could we save some money by perhaps turning off the servers at night and on weekends? Our costs are currently very low and can’t be reduced without fully cutting off access, but the investigation might reveal something about what it is like to run a campus Generative AI service.
I’ve a few sources of historical data to work with. Each gives a partial view, but there’s pretty strong consensus in what they say.
First, I can look at the last 90 days of message traffic metadata on BayLeaf Chat. I can’t look further back than 90 days because, per our data retention policy, that’s literally all we store. So, this’ll tell us about the tail end of spring quarter and some quieter and less structured usage in the summer quarter.
Next, I have six months of logs from the OpenCode coding agent harness on my laptop and more than two years of activity on a friends-and-family instance of Open WebUI. Together, these represent something close to what we might expect for the BayLeaf API and BayLeaf Chat services over a longer period of time, but biased toward my personal usage patterns.
Before I launch into what I found, think about what you might expect the data to show. If UC Santa Cruz were a stereotypical business organization, you might expect steady weekday usage during business hours, and maybe even a clear lunchtime notch. Maybe Monday mornings would show users accessing the system to get oriented for the week, and Friday afternoons would show a gentle ramp down as people left early. Now, consider that we’re at a university, with folks staying up late to prepare for exams, balancing social lives, perhaps specifically *not* accessing the services while they are in scheduled lectures. What’s that usage going to look like?
According to the data, weekends aren’t a big thing. Whatever effects you might think would push weekday and weekend patterns apart seem to mostly cancel out. Usage on Saturdays and Sundays falls between the highest and lowest weekday usage in all three datasets. The location and ratio of peaks and troughs vary a bit by dataset and by whether you aggregate by session count, message count within a session, token count within a message, and so on. My personal data shows I’m almost twice as active on Thursdays as I am on Fridays, but the campus data suggests this is just a personal quirk. Someone else with a different schedule might have a consistently different profile.
Time of day is where things get a bit more interesting. Across all three datasets, there’s a clear on-off pattern, with activity starting around 8am and turning off around midnight. In the off period, we’re not fully off. Activity is minimal but nonzero (even in my personal datasets) at 4am. Meanwhile, 4pm is a daytime peak. My personal data shows distinct notches at 5pm, for commuting home, and 7pm, for cooking and eating dinner. However, these notches only appear when you look at user message counts. For requests at the API level (which include agents following up on tool call results, etc.), they go away. I often leave my agents cooking up some code while I’m cooking my family dinner.
Overall, there’s a wake-sleep pattern that comes from most of the UC Santa Cruz campus community being in the same place at the same time. We’re maybe 20 times as active in wake mode as in sleep mode. But is this pattern exploitable?
By analogy to electrical power generation, it can be cheaper and more efficient to keep some kinds of generators running continuously around the clock. To handle demand spikes, you might then bring peaker plants online at critical moments. This isn’t even generally true across the full range of generation technologies, but let’s see how far the analogy gets us.
If we ran BayLeaf using on-campus GPUs (we don’t), maybe we could turn off those servers at midnight and wake them up at 8am. You’d hope that, if we had 20 GPUs, we’d simply turn off 19 for sleep mode so that the service remained available to night owls. However, the situation isn’t so simple. For a given size of model, there’s a minimum hardware configuration needed to generate tokens at a reasonable rate. We outsource inference to third-party services because, even at our 4pm peak usage, we can’t fully utilize a single dedicated inference server.
If our wake-sleep patterns were 12 hours on, 12 hours off, perhaps we could coordinate with a sister campus in (*checks globe*) Iran, Kazakhstan, or western Russia. They’d be active when we weren’t. With a 16-hour active period, however, we’d have overlap. Maybe we could band together with other West Coast universities (perhaps in some sort of system of universities in California) and reach enough aggregate demand that diurnal resource scaling starts to make sense. Until then, it seems prudent to pool our usage with the much larger national market and let providers with much more comprehensive data decide how many servers to keep active at various times of day.
I realized I’ve gotten to the end of this data-oriented post without sharing any concrete data. So, uh, here’s a snapshot of my personal coding agent data, which turns out to be fairly representative of campus-wide usage as well:


