tweetindex
EN

Shobhit Banga ✓

@shobhitbanga

Scorekeeper for voice AI. Building independent, real-world evals to help labs build the best voice models. Co-founder @voicearena_ai | prev @JoshTalksAI

3 130Followers
562Following
1 901Posts total
581.2KViews on collected posts

Latest posts

@shobhitbanga @GoogleDeepMind A real-human, multi-turn benchmark is the right direction. I would publish the task mix, interruption points, rater protocol, and failure cases too. A #2 result becomes much more useful when readers can see where Gemini 3.8 Live wins, stalls, or reco 5 views · 0 likes · 0 reposts · 0 replies Open on X →
@shobhitbanga @GoogleDeepMind @OfficialLoganK @GoogleDeepMind y'all cooked 🔥 congrats on awesome results on the @voicearena_ai leaderboard 32 views · 0 likes · 0 reposts · 0 replies Open on X →
@shobhitbanga @GoogleDeepMind If you ever publish the stability metrics, I'd love to compare notes on my own dashboards. 65 views · 0 likes · 0 reposts · 0 replies Open on X →
@shobhitbanga @voicearena_ai Okay, this is seriously impressive. 🔥 A benchmark where humans judge how human the AI actually feels 3.2K views · 4 likes · 0 reposts · 0 replies Open on X →
On Task Completion, though, it's a different story, the gap is much closer. Votes are still coming in and we will keep updating the leaderboard. This is a work in progress. Tag the model labs you would like to see added to this benchmark. https://t.co/HPGJT02Uec 447 views · 5 likes · 0 reposts · 3 replies Open on X →
Here's one of those calls: gpt-realtime, our top model on Humanness, next to a real caller, on a gift-budget question. Same conversation, side by side. Play it and see how far 18% actually sounds. https://t.co/jrTxizv9dX
0:34
491 views · 1 likes · 0 reposts · 1 replies Open on X →
Read every pair head to head and the same thing holds. The human agent is favoured against every single model: 82% against gpt-realtime, and up to 87% against Grok Voice and Inworld. The models sit within 60-40 of each other, so the race between them is close. The race against
1.2K views · 3 likes · 0 reposts · 1 replies Open on X →
The first board is Naturalness: how much each call sounds like a real human conversation. The human agent sits at 1244 Elo. Every model is a wide gap below the human: › gpt-realtime 984 (-260) › Gemini 3.1 Flash Live 970 (-274) › GPT Live 1 967 (-277) › Grok Voice 920 https://t
12.2K views · 8 likes · 0 reposts · 4 replies Open on X →
Introducing Jarvis Bench v0.5, @voicearena_ai's conversational agent benchmark. We've been obsessed with one question at VoiceArena: why do voice agent demos sound incredible, benchmarks say models are near-perfect, and yet you probably didn't have a single real conversation htt
4:15
563.5K views · 149 likes · 11 reposts · 30 replies Open on X →

Against accounts of the same size

9 posts from the last 90 days, next to the under 10K follower range. below its peers on both reach and engagement.

Median views491this account4 175median for under 10K
Reach, %15.69%this account250.82%median for under 10K
Engagement, %0.10%this account1.41%median for under 10K
MetricThis accountMedian for under 10KRatio
Median views per post4914 1750.12×
Reach (views ÷ followers)15.69%2.5× audience0.06×
Engagement rate0.10%1.41%0.07×

Others in this range →   Compare with another account →   How these benchmarks are built →

Growth & engagement

How the posts we collected actually performed: views and reaction rate post by post, what the audience did with them, and where the follower count goes.

Views per post

563.5K15 Sep
12.2K
1.2K
491
447
3.2K
6517 Sep
3218 Sep
5

Last 9 collected posts, oldest on the left. The scale is logarithmic: one post can outrun the rest a hundred times over.

Engagement rate per post

0.04%15 Sep
0.10%
0.34%
0.41%
1.79%
0.13%
0.00%17 Sep
0.00%18 Sep
0.00%

Reactions — likes, reposts, replies and quotes — divided by views. Median for under 10K accounts is 1.41%.

What the audience does

Likes46.7%170 in total
Reposts3.0%11 in total
Replies10.7%39 in total
Quotes17.0%62 in total
Bookmarks22.5%82 in total

Share of every reaction we collected for this account. Replies mean argument, reposts mean endorsement, bookmarks mean the post was worth keeping.

The follower curve appears once this account has two daily snapshots — we take one a day, and this one is on its first.

Similar accounts