Our first impression with OpenAI's GPT-OSS models

Discover how Q2BSTUDIO uses OpenAI's GPT-OSS-20B and GPT-OSS-120B models to offer custom software solutions, enhance processes with artificial intelligence, and ensure security and privacy in cybersecurity. Contact us to evaluate and integrate these technologies into your business projects.

sábado, 16 de agosto de 2025 • 3 min read • Q2BSTUDIO Team

Artificial-Intelligence-

At Q2BSTUDIO, a custom software and application development company specializing in artificial intelligence and cybersecurity, we have integrated OpenAI's GPT-OSS-20B and GPT-OSS-120B models into our ForgeCode tool to evaluate their performance in local environments and real-world workflows.

We have translated our first impressions and results into this practical summary for teams seeking custom software solutions, custom applications, AWS and Azure cloud services, business intelligence services, and AI implementations for businesses.

Integration and privacy: being open-weight models under the Apache 2.0 license, GPT-OSS models allow running inference locally, maintaining total code privacy, and adjusting prompts without sending data to the cloud. This is key for companies concerned about cybersecurity and intellectual property protection.

Performance in benchmarks: in our tests, GPT-OSS-120B showed very competitive scores in reasoning and mathematical tests comparable to proprietary models. MMLU close to 90.0 versus a 93.4 reference, GPQA Diamond around 80.1 versus 83.3, solid results in AIME-style mathematical challenges with 96.6 in 2024 and 97.9 in 2025. The 20B model surprises with its efficiency and accuracy for lightweight tasks, with scores around 85.3 in MMLU and 96.0 in AIME 2024.

Speed and terminal experience: in ForgeCode, the models provide responses in less than a second in many cases, even with prompts involving multiple files or phases. This enables real-time CLI assistance for developers, from commits to refactorings or schema migrations.

Accuracy with commands and tools: we observed high accuracy when generating git commit messages, TypeScript interface scaffolding, and in tasks integrated with external tools. The combination of AI agents and the ability to run locally facilitates more reliable and auditable workflows.

Behavior in multi-stage actions: some cases show intermediate stops in responses, for example stopping at phrases like "Phase 1 begins here" without completing. We are fine-tuning prompts and follow-up strategies to improve follow-through in multi-stage tasks.

Transparency and optimization: the open-weight nature allows benchmarking, fine-tuning, and openly sharing results. That level of transparency drives innovation in the ecosystem and helps providers and internal teams optimize models according to their needs.

Choosing the model according to the task: for lightweight edits and quick responses, we recommend GPT-OSS-20B. For deep reasoning over large codebases, we prefer GPT-OSS-120B. In ForgeCode, switching models is immediate and simple from the command-line interface.

Key benefits for businesses: privacy and control over data and code, performance and speed for real-time development assistance, transparency for auditing and continuous improvement, and a boost to innovation in enterprise artificial intelligence solutions. These benefits align with services we offer at Q2BSTUDIO such as custom software development, artificial intelligence solutions, AI agents, cybersecurity services, AWS and Azure cloud service implementations, business intelligence solutions, and Power BI dashboards.

Recommended use cases: CI/CD integrations with automatic change validation, code generation and refactoring in large repositories, specialized assistants for technical support teams, business intelligence pipelines combining local models with Power BI visualizations, and AI agents to automate repetitive tasks while maintaining compliance and security.

How to try it: you can run the models locally with ForgeCode and test GPT-OSS-20B or GPT-OSS-120B from the public repositories on HuggingFace, or follow ForgeCode's quick access path to get started in your terminal. Testing locally makes it easier to evaluate real performance on your workloads and allows Q2BSTUDIO to help with integration, customization, and ensuring the solution meets security and scalability requirements.

At Q2BSTUDIO, we offer support throughout the entire cycle: from consulting to select the right architecture, custom software development and custom applications, to deployments on AWS and Azure cloud services, cybersecurity audits, AI agent integration, and business intelligence projects with Power BI. Our goal is for companies to harness the power of artificial intelligence without sacrificing privacy or control.

Conclusion: the arrival of OpenAI's GPT-OSS-20B and GPT-OSS-120B models represents an important step toward powerful, local, and transparent AI. At Q2BSTUDIO, we are already applying these capabilities to offer custom software solutions, enhance processes with AI for businesses, and ensure security with good cybersecurity practices. If you want us to help you evaluate or integrate these technologies into your projects, contact us and we will explore the best personalized solution for your business.

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.