In their paper *Frugal GPT*, Lingjiao Chen, Matei Zaharia, and James Zou describe three ways to cut the cost of calling a large language model: prompt adaptation, LLM approximation, and LLM cascade. The cascade tries cheaper models first and can match the best individual model with up to 98 percent cost reduction on their tasks. I built a small agent that checks apartment listings against a renter's requirements, such as a two bedroom in Austin under $1,500 that allows dogs. My first version sent every listing to a strong model, resulting in 2,500 model calls to find 101 real matches.
I tested three starting points: sending every pair to the model, searching by city, and running a database query on city, bedrooms, and rent before any model call. Using a fixed sample of 500 listings from a public dataset of 2019 United States rental ads, I scored every version against five renter requirements. The strong model was Open AI's gpt-6-sol (Sol), and the cheaper model was gpt-6-luna (Luna). I traced every call with Weights & Biases (W&B) Weave.
Against the database query starting point, the final version cost about 25 times less—$0.008 versus $0.20 for the same 2,500 checks—and returned all 101 matches with no wrong ones. Letting code settle the pets field made up about 44 percent of the cost drop, and using the cheaper model with a strong backup made up about 39 percent. A shorter prompt and reused answers made up the rest.
Three findings surprised me: Weave's cost column showed $0.0000 for every call; Sol missed a match 10 times out of 10 with the full listing but found it 10 times out of 10 with only the title and body; and both models reported a confidence of 5 on almost every answer, so a rule that escalates below 5 almost never fires. Only one of the 101 matches depends on reading text, so the test measures cost far better than understanding.
Source: Towards Data Science · Summarized by HeadlinesBriefing