CalculatedContent/WeightWatcher
quality grade C, 51 out of 100The WeightWatcher tool for predicting the accuracy of Deep Neural Networks
- stars
- 1.8k
- stars gained this week
- +2
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Benchmarks, leaderboards, LLM-as-judge harnesses, testing and red-teaming.
Signals: benchmark, evaluation, llm-evaluation, leaderboard, llm-eval, red-teaming, testing, model-evaluation
951 results
The WeightWatcher tool for predicting the accuracy of Deep Neural Networks
A toolbox to iNNvestigate neural networks' predictions!
No description
Readymade evaluators for agent trajectories
a high performance library for building cache simulators
List of papers on hallucination detection in LLMs.
UpTrain is an open-source unified platform to evaluate and improve Generative AI applications. We provide grades for 20+ preconfigured checks (covering language, code, embedding use-cases), perform root cause analysis on failure cases and give insights on how to resolve them.
AI Red Teaming playground labs to run AI Red Teaming trainings including infrastructure.
C2concealer is a command line tool that generates randomized C2 malleable profiles for use in Cobalt Strike.
A Huge Learning Resources with Labs For Offensive Security Players
Cover your tracks during Linux Exploitation by leaving zero traces on system logs and filesystem timestamps.
BigBountyRecon tool utilises 58 different techniques using various Google dorks and open source tools to expedite the process of initial reconnaissance on the target organisation.
Template-Driven AV/EDR Evasion Framework
Tips and Tutorials for Bug Bounty and also Penetration Tests.
A Windows reverse shell payload generator and handler that abuses the http(s) protocol to establish a beacon-like reverse shell.
一个攻防知识库。A knowledge base for red teaming and offensive security.
Complete Mandiant Offensive VM (Commando VM), a fully customizable Windows-based pentesting virtual machine distribution. [email protected]
ROS package for the Perception (Sensor Processing, Detection, Tracking and Evaluation) of the KITTI Vision Benchmark Suite
An automatic evaluator for instruction-following language models. Human-validated, high-quality, cheap, and fast.
Framework to evaluate peformance of ROS 2
Cypress authentication flows using social network providers
Examples how to test Alembic migrations for Yandex Backend Development school & Moscow Python meetup №69
A pytest plugin for preserving test isolation in Flask-SQLAlchemy using database transactions.
👓 A collection of all the Google developer and engineering blogs related to programming, security, opensource, testing, android, youtube, etc.
24,537 repositories in the index in total.