The RL vs SFT Management War: How Moonshot AI's Yang Zhilin Is Redefining Crypto Team Dynamics
From the ashes of 2017, when ICO whitepapers promised the moon yet delivered empty binaries, I've learned to read between the lines. So when I stumbled upon a thread by Moonshot AI’s founder Yang Zhilin comparing team management to reinforcement learning (RL) and supervised fine-tuning (SFT), I didn’t just see a clever analogy—I saw a potential blueprint for the next phase of crypto organizations. The timing couldn’t be more critical: as the crypto industry shifts from speculative narratives to sustainable infrastructure, the way we structure our teams will determine which protocols survive the coming wave of institutional scrutiny.
Yang’s core thesis is deceptively simple. SFT, in AI training, is the process of feeding a model labeled examples—essentially telling it exactly what to output. In management, this translates to micromanagement: step-by-step instructions, rigid KPIs, and hierarchical command chains. RL, by contrast, lets the model explore its environment, receiving rewards for actions that lead to desired outcomes. Applied to human teams, it means setting clear goals and incentives, then giving members the autonomy to find their own path. Yang claims Moonshot AI operates primarily on RL principles, with SFT used only for edge cases like security red lines.
As someone who spent years auditing smart contracts and analyzing tokenomics, I immediately recognized the framework. In crypto, we’ve seen both extremes. The early DAOs of 2016 were pure RL: chaotic, permissionless, and prone to reward hacking—remember The DAO hack? The subsequent era of centralized exchanges and venture-backed projects leaned heavily SFT: strict order books, mandated listing criteria, and top-down governance. Neither survived intact. Yang’s insight is that the healthiest organizations need a dynamic balance, not a static choice.
But let’s apply the rigor that the analogy deserves. In AI training, RL suffers from three notorious failure modes: sparse rewards, credit assignment, and reward hacking. In a corporate RL setting, these map directly to employee burnout, inability to attribute success to specific actions, and gaming the system. Yang acknowledges the last one—he warns that “complete RL can lead to gaming the system.” Yet his analogy stops short of addressing how to build a robust reward function. In my experience analyzing the 2022 Terra collapse, the “Anchor protocol” offered a fixed 20% yield (a reward function) that gamed itself: the protocol bled reserves until death. The same happens in teams when bonuses are tied to vanity metrics like code commits or hours billed. Without a mechanism for continuous reward recalibration—something AI researchers call “reward shaping”—the RL approach becomes a ticking bomb.
From the ashes of 2017 to the fluidity of DeFi, I’ve watched narrative cycles repeat. During DeFi Summer in 2020, I manually tracked 50+ yield farming strategies and interviewed founders. The most successful projects didn’t just have great TVL; they had teams that operated like RL agents: testing hypotheses, iterating fast, and sharing a collective curiosity. Uniswap’s early team, for example, had no strict product roadmap—they simply defined “increasing liquidity efficiency” as the goal and let experiments lead the way. That ethos produced innovations like concentrated liquidity and automated range orders. In contrast, projects with heavy SFT management—rigid hierarchies pushing predefined features—often shipped technically sound products that nobody used.
But here’s the contrarian angle that Yang’s narrative omits: RL management scales poorly beyond elite, homogeneous teams. In AI, training a single RL agent is hard enough; training a fleet of agents with shared goals is exponentially harder. The crypto industry has already experienced this scaling problem in DAO governance: small, tight-knit DAOs like Moloch work beautifully, but larger ones like MakerDAO require extensive rulebooks and bureaucratic oversight—essentially SFT injections. As Moonshot AI grows from 50 to 500 employees, its RL culture will face an alignment crisis. Who decides the reward function when multiple teams have conflicting priorities? How do you prevent one team’s “exploration” from breaking another’s production system?
Moreover, the persona’s core opinions offer a sharp critique: the “blue chip NFT label is a trap” because when liquidity dries up, floor prices prove nothing. The same goes for management labels. A company that proudly calls itself “RL-first” may attract idealists but alienate essential risk-averse talent—compliance officers, legal counsel, community managers. In the AI industry, these roles are the equivalent of SFT: they ensure the model behaves safely. Without them, RL organizations become vulnerable to the same pitfalls as unwrapped smart contracts: one exploit and the whole system implodes.
I recall auditing a DeFi protocol in 2021 where the team had a strict “no meetings” policy (an RL extreme). Developers thrived in autonomy, but the lack of aligned reward signals led to duplicated work and a critical vulnerability that went unnoticed for weeks. In contrast, a larger competitor using a hybrid model—weekly SFT-style stand-ups with RL-style project ownership—caught similar bugs early. The lesson: the optimal strategy isn’t pure RL or SFT, but a context-dependent blend. Yang’s own work on Kimi, with its long-context window and agentic tools, suggests Moonshot AI already understands this: the model combines RL-based exploration with SFT-based safety filters.
From the ashes of 2017 to the fluidity of DeFi, we’ve learned that narratives alone don’t build value—execution does. Yang’s management philosophy is a fascinating narrative, but it needs to be tested against the cold reality of on-chain metrics. Will Moonshot AI’s employee retention outperform teams using traditional structures? Are their product iteration cycles actually shorter? Until we see data, we must view this as a compelling hypothesis, not a proven framework.
So where does this leave the crypto industry? The next wave of innovation will likely come from projects that institutionalize the RL/SFT hybrid. We’re already seeing glimpses: Arbitrum’s “R&D DAO” that allocates grants with minimal oversight (RL) but requires quarterly reports (SFT). The real prize is designing reward functions that automatically adjust to market conditions—a kind of “decentralized management” that aligns individual incentives with protocol health.
As for Yang Zhilin and Moonshot AI, they have the chance to become the case study that validates or refutes the RL management narrative. But like any good crypto thesis, we should follow the data. Is the reward system transparent? Are employees being rewarded for outcomes they genuinely control? Or is the company’s narrative just another form of liquidity mining—exciting, but ultimately inflationary?
The narrative is shifting, and this time it’s not about tokens or TVL. It’s about the code that runs our teams. From the ashes of 2017 to the fluidity of DeFi, the question remains: can we build organizations that learn as fast as the protocols they create?