r/GeminiAI • u/tru3_gentleman • 4h ago
Help/question POST your results or I'm calling AGY useless
Genuinely asking because at this point I’m wondering if I’m the problem lol. With Codex and Claude I’m getting like a 7/10 average on first try — decent function, decent design, and if something breaks, they actually troubleshoot it.
Then Gemini just... doesn’t understand what I’m asking half the time, and when something isn’t working it’ll literally make up explanations or pretend a feature/tool exists and that everything is fine.
Like bro, I’m looking at the screen. 💀
Same prompts, same kind of tasks, completely different results.
So WHAT THE HECK is Google doing? Is Antigravity the problem, Gemini itself, or am I somehow using it completely wrong?
Because the results people post here vs. my actual experience feel like two completely different products.
•
u/Jarr11 2h ago
I have this issue alot. Codex just gets it and performs great almost all of the time, but the quota doesnt last long enough. But my Gemini quota I can't seem to burn through quick enough, however because the model doesnt spend enough time reasoning, quite often it just gets things wrong.
If you want a capable model, use Codex/Claude. If you want a model with near-unlimited quota, use Gemini. But you cant have both capability and enough quota 😆
•
u/reinka 2h ago
Been having the same issue that it completely fabricates things. Yesterday I asked it to research potential managed postgres DB Services for my project and it made up a pricing table and feature list of an actual provider. I asked Astra to check on the proposal and during verification it spotted the invented parts and claims. When I asked Gemini 3.8 flash about it it responded something along the lines that the feedback is totally right and he himself was relying to much on the marketing fluff etc. on the website. Dude, I literally have a skill that's supposed to guard against it and asks to verify assumptions and cut the fluff...
I don't even let it do complex coding anymore because too often Sol had to fix its mess: it seems like it doesn't have the depth to scan through the code base and see the bigger picture even though I specifically prompt for it and have specific skills.
My last attempt will be Astra as an orchestrator and Gemini 3.8 flash as the implementer, supervised and reviewed by Astra. Haven't had the time to test it yet but if that turns out to be another failure I'll just give up on Gemini until they release another Pro model.
I really don't understand how it performs so well on the DeepSWE benchmark. I guess I have to look up those prompt styles and ask Astra to delegate work to Gemini the same way...
•
u/tru3_gentleman 1h ago
Yep, same here, thanks for your answer, been thinking is it only me or is it really a scam...
•
u/AutoModerator 4h ago
Hey there,
This post seems feedback-related. If so, you might want to post it in r/GeminiFeedback, where rants, vents, and support discussions are welcome.
For r/GeminiAI, feedback needs to follow Rule #9 and include explanations and examples. If this doesn’t apply to your post, you can ignore this message.
Thanks!
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
•
u/Exotic_Fig_4604 3h ago
I cannot post my results from Claude, because it runs out of token after building half a feature.
Does that answer your question?
Claude is absolutely unusable for me, because it just runs out of tokens every 5 minutes.
If you find a million Dollars lying around to build a hello world app, let me know.
That being said, Antigravity is definitely not the problem, Gemini is. Antigravity is actually my favourite agentic IDE, except for their crappy models.
•
u/tru3_gentleman 3h ago
If it is not claude and gemini, what are you using then? o_O
•
u/Exotic_Fig_4604 3h ago
This week I started using Muse Spark 1.3
Its far from perfect, but its dirt cheap and reasonably fast. Definitely better performance than the current Gemini models.
I am using it with OpenRouter in Intellij, where I am sadly lacking all the QoL features of AGY, but even with that its better.
I will look for a better harness on the weekend.
•
u/aero_sock 2h ago
I don't know, I have had pretty positive experiences with antigravity. I run flash 3.8 on high 99% of the time. I gave Claude a shot, and like yeah ig it's smarter, but the difference isn't that big, and Claude pro is 4x the price of gemini pro student(20usd vs 5usd) and afaik the effective limits are much smaller than in agy. I tried it in both c++ embded projects as week as webdev and it gets the job done for me.
•
u/DropEng 4h ago
Are you looking for this? https://deepmind.google/models/model-cards/gemini-3-8-flash/
•
u/blank_detention 4h ago
That model card link is not the flex they think it is, just a shiny brochure for the same broken toy.
•
u/Then_Bake_6524 3h ago
what people can't really separate is that most of the benchmarks show the ability of the model to do what the user asked for, even the things the user didn't think of (such as debugging, extra features, QoL, etc).
examples are GPT-6 Astra building you a carbon copy of a certain app/game with the same things and minimal bugs, or actually doing anything you ask it to, like playing a game. while Gemini gives up way too early and does the bare minimum (yes, i saw all the posts showing 1980s-like web game graphics).
basically, what most regular users care about is the ability to do something with only a bare "hey, i am bored today, build me X thing." and the model doing it with no questions or mistakes.
yes, Gemini works perfectly with easy tasks, any model works perfectly with easy tasks.