Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Buckets new
  • Docs
  • Enterprise
  • Pricing
    • Website
      • Tasks
      • HuggingChat
      • Collections
      • Languages
      • Organizations
    • Community
      • Blog
      • Posts
      • Daily Papers
      • Hardware
      • Learn
      • Discord
      • Forum
      • GitHub
    • Solutions
      • Team & Enterprise
      • Hugging Face PRO
      • Enterprise Support
      • Inference Providers
      • Inference Endpoints
      • Storage Buckets

  • Log In
  • Sign Up

RewardHacking

Activity Feed

AI & ML interests

None defined yet.

Recent Activity

xinpeng  submitted a paper 4 days ago
OPD Before RL: Warm-Starting Rubric-Based RL with On-Policy Distillation
tongliuphysics  authored a paper 12 months ago
Temperature-scaling surprisal estimates improve fit to human reading times -- but does it do so for the "right reasons"?
tongliuphysics  authored a paper 12 months ago
FocalPO: Enhancing Preference Optimizing by Focusing on Correct Preference Rankings
View all activity

wang's profile picture Tong Liu's profile picture

rewardhacking 's datasets

None public yet
Company
TOS Privacy About Careers
Website
Models Datasets Spaces Pricing Docs