Practice / selecting and filtering

Tip-Top Baristas

medium

Corporate at The Daily Grind is launching a "Tip-Top Barista" award, and naturally they want data instead of vibes. The metric: which orders tipped at least 20% of the drink price?

Same orders DataFrame as before:

Your task:

  1. Compute each order's tip percentage as tip / price, in a column named tip_pct.
  2. Keep only orders where tip_pct is at least 0.20.
  3. Return order_id, drink, and tip_pct.

Assign the DataFrame to result. Row order doesn't matter, and tip_pct is compared with a float tolerance — no need to round.

ordersinput DataFrame

Schema
columntype
order_idlong
drinkstring
sizestring
pricedouble
tipdouble
Sample rows (first few)
order_iddrinksizepricetip
1001oat lattelarge6.51
1002drip coffeesmall30
1003caramel macchiatolarge7.252
1004flat whitemedium51.5
1005cold brewlarge5.50.5
1006espressosmall2.50.75
1007mochalarge6.750.25
1008lattemedium5.252
1009gold-dust affogatosmall6.50.75
1010barrel-aged cold brewlarge60.5
Expected output shape
order_iddrinktip_pct· 5 rows
Hint

Build the ratio with withColumn("tip_pct", F.col("tip") / F.col("price")), then filter on the new column. You can reference a column you just created in the very next chained step — transformations stack.

Lesson refresher

This problem builds on Selecting & Filtering (~6 min). Pop it open in a new tab if you want a quick recap.

Loading editor…
Hit Run to execute your code and see the output here.