Practice / aggregations

Barista Averages

medium

Corporate at The Daily Grind wants to spotlight its busiest stores — but "busy" means real money, not a high sticker price on one lucky order.

You have the orders DataFrame:

Your task: for each store, compute the average order price as avg_price, but return only stores whose total revenue is greater than 15.00. Return store and avg_price.

Assign your answer DataFrame to result. Row order doesn't matter, and avg_price is graded with a small float tolerance — no need to round.

ordersinput DataFrame

Schema
columntype
order_idlong
storestring
pricedouble
Sample rows
order_idstoreprice
1Downtown5
2Downtown6
3Downtown5.5
4Airport4
5Airport3
6Pier7
7Pier9
8Harbor15
Expected output shape
storeavg_price· 2 rows
Hint

Aggregate two things at once: F.avg("price").alias("avg_price") and F.sum("price").alias("total"). Then filter on the total, and finally select just store and avg_price. HAVING is just a filter after the agg.

Lesson refresher

This problem builds on Aggregations & groupBy (~7 min). Pop it open in a new tab if you want a quick recap.

Loading editor…
Hit Run to execute your code and see the output here.