article
Choosing a framework for a project is not obvious. Although surveys, expert panels, and targeted benchmarks provide guidance, those approaches are costly and difficult to implement. Moreover, they often fail to synchronise with evolving trends. To overcome this limitation, the MultiEval approach automates the evaluation protocol by leveraging six large language models (LLMs): GPT-5, Claude 4.5 Sonnet, Gemini 2.5 Pro, DeepSeek-V3.2, Qwen3-Max, and Grok 3. The approach uses a shared set of standardised prompts to generate ratings on a scale of 1 to 5 across six key dimensions: adoption, community support, learning curve, performance, security, and maturity. We then correlate those scores against GitHub and Stack Overflow signals to assess how well they mirror real-world ecosystem data. For backend frameworks, the six-model ensemble aligns with expert judgments 92% of the time ($\kappa=0.85$ [95% BCI: 0.81-0.88]) and outperforms every single model. Frontend scores exhibit higher variance because developer experience is inherently subjective, yet they remain strongly correlated with public ecosystem data ($r=0.98$ [95% BCI: 0.95-0.99]). A retrospective backtest that withheld all post-2020 information demonstrates the ensemble’s ability to identify long-term winners early (78% top-3 accuracy [95% BCI: 68-86 %]). Taken together, a heterogeneous model ensemble combined with a fixed scoring protocol offers a fast, repeatable method for rating software frameworks at scale.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1109/iraset68627.2026.11538764
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.