Guide Labs Releases Open-Source Interpretable AI Model Steerling-8B
View original sourceGuide Labs, a San Francisco-based startup founded by CEO Julius Adebayo and chief science officer Aya Abdelsalam Ismail, has developed an open-source large language model (LLM) called Steerling-8B. This model, released on Monday, addresses the challenge of interpretability in deep learning models by tracing each token to its training data.
- Unlike traditional models that struggle with complexities like humor or politics, Steerling-8B uses a unique architecture with a 'concept layer' that categorizes data into traceable entities, improving interpretability.
- Adebayo previously co-authored a 2018 paper highlighting inadequacies in existing methods of understanding deep learning models, inspiring the new approach.
- The company claims Steerling-8B can achieve 90% performance of current frontier models while requiring less data.
This new architecture is poised to benefit consumer-facing applications by allowing better control over sensitive content and compliance with regulations, such as in finance or scientific research. Concerns about the loss of emergent behaviors are addressed as Steerling-8B can still discover new concepts independently.
Guide Labs, backed by a $9 million seed round from Initialized Capital and emerging from Y Combinator, plans to further build and offer API access to this model. Adebayo emphasizes the importance of manageable AI for future societal integration.