The recent release of Smaug-Llama-3-70B-Instruct has generated a lot of buzz, but also some misconceptions that need addressing. As the lead developer, I want to clarify a few key points to ensure everyone has accurate information. First and foremost, contrary to some claims, we did not train the model on benchmark datasets. The training datasets used are clearly listed on the model card, including Orca-Math-Word, CodeFeedback, and AquaRat. These datasets were chosen for their robustness and relevance to real-world applications.

Another point of contention is the validity of the benchmarks we used. While some argue that benchmarks like MT-Bench are maxed out and less compelling, we selected MT-Bench and Arena-Hard because they correlate well with general real-world usage. Arena-Hard, in particular, was designed to align closely with the Human Arena leaderboard, making it a reliable indicator of performance. Although we haven't been able to get our model on the Human Arena leaderboard yet, we are actively trying and hope to succeed soon.

Lastly, it's important to address the skepticism around the sensationalist nature of some promotional content. While the style of certain Twitter accounts may not appeal to everyone, our commitment to scientific rigor remains unwavering. Our goal is to create the best open-source LLM possible, bridging the gap with closed-source models. My background in AI, including my tenure at DeepMind, underscores our dedication to this mission. For those seeking a more neutral tone, I invite you to follow my new Twitter account for updates and insights.

In summary, Smaug-Llama-3-70B-Instruct is built on a foundation of transparency and scientific integrity. We welcome further questions and discussions to clear up any remaining doubts and to continue improving the model for the community.