HeadlinesBriefing favicon HeadlinesBriefing.com

GRPO se laghu language models ko sikhata hai

Towards Data Science •
×

Ek language model 'let me double-check that' likh sakta hai aur phir bhi multiplication galat ho sakti hai. Yeh bhi likh sakta hai 'wait, let me reconsider,' aur galt answer de sakta hai. Koi bhi yeh sentence sochne ka proof nahi hai.

Agar goal problem solve karna hai, toh sirf akhir number important hai. Deep Seek's R1-Zero ne reinforcement learning ko mere se pehle kisi supervised fine-tuning stage ke bina attention diya tha. Iske researchers ne behaviors jaise approach ko revisit karna aur difficult problems par zyada tokens use karna report kiye the.

Deep Seek-R1, in contrast, ek broader training pipeline use kiya tha, jo cold-start data include karta tha, toh humein do training processes ko interchangeable nahi samajhna chahiye. Jo mujhe sabse zyada interest karta hai wo hai Group Relative Policy Optimization, ya GRPO jo hum logon ko pata hai. Basic setup ko bohot kam feedback chahiye hota hai.

Model ek sawal par koi baar-baar generate kar sakta hai, har baar score milega, aur behavior adjust karne ke liye farak use kar sakta hai. Humein even final answer check karne ke liye us solution ko likhna zaroori nahi jiska wo mimic kare. Memory reduction techniques ke saath combine karne par, yeh small reasoning experiments ko accessible banata hai.

Do cheezein hain jo model ya GPU choose karne se pehle samajhne lagegi: score guide karna kaise update karta hai, aur jab scoring rule galat cheez ko reward deta hai toh kya hota hai. Ek simple arithmetic example do baato explain karta hai. Chaar possible responses ek sawal ke liye aur kaise reward function har ek ko score deta hai.

Ek problem jis answer ko check kiya ja sakta hai: Ek dukan ke paas 6 boxes hain, har box mein 8 items hain. 6 items bech di gayi. Kitne items bache? Jawab 42 hai. Ek Python expression isko verify kar sakti hai.

Ab, yeh simple check humein koi valuable deta hai: final answer score karne ka independent tarika. Humein doosre language model ki zaroorat nahi padti ki yeh decide kare ki response sunne mein convincing lag raha hai ya nahi. Training dataset mein, sawal model ko jayega jabki expected answer verifier ke saath rahega.

Model response generate karega, verifier score karega, aur training algorithm model behavior adjust karne ke liye score use karega. Asaan. Isliye humein supervised targets ke taur par solved solutions dene bina question-and-answer pairs use kar sakti hain.

Reinforcement learning with verifiable rewards ka appeal yeh hai ki feedback outcome check se aayega. Real inventory system ke liye, ordinary arithmetic still sensible solution hai. Language model expensive calculator ke saath opinions hoga.

Yahan, example seekhane mechanism ko easy inspect karne mein madad karta hai. Yeh bhi dikhata hai ki hamare verifier kahan gir gaya hai, kyunki final integer check karna yeh nahi batata ki uske peeche ka reasoning kaise maansik hai. GRPO Deep Seek Math work mein conventional PPO-based training ke alternative ke roop introduce kiya gaya tha, memory requirements ko solve karte hue.

Common actor-critic implementations of PPO train a value estimator, often called a critic, alongside the policy. That estimator supplies a baseline for judging an action’s outcome. GRPO apne baseline ko same prompt ke liye sampled responses ke group se obtain karta hai.

Sirf doosre critic se poochne ke baare mein ki model kaise karna chahiye, it compares how well several attempts actually did. Imagine model yeh char answers produce karta hai: Attempt Final answer Correctness reward 1 42 1, 2 24 0, 3 42 1, 4 48 0. Group ki average reward 0.5 hai.

Attempts 1 aur 3 us average se behtar performance diye; attempts 2 aur 4 se behtar nahi. Original outcome-reward formulation standardizes that comparison: [Ai=ri−rˉσr+ϵ]. Yeh dikhata hai ki attempts 1 aur 3 ka standardized reward 0 se upar hai, jabki attempts 2 aur 4 ka standardized reward 0 se neeche hai.

Is tarah, GRPO ek group of responses ke comparison ka use karta hai model update guide karne ke liye, bina koi external critic ke.