Item detail
github.com

InternScience/Agents-A1

InternScience/Agents-A1 is a model release that RepoRadar is tracking in its New Models section, currently rated Gold tier with a 'try now' verdict. Its strongest signal is workflow potential, scored 9.2 out of 10.

Score8.1
Popularity1.0
Risklow
TierGold
Score breakdown
Usefulness9.0
Novelty8.0
Momentum8.0
Maturity6.4
Open-source/build8.4
Evidence7.2
Workflow potential9.2
Setup ease6.4

Popularity is tracked separately. Support, ads, sponsorships, and tips never affect these signals.

Why it matters

Useful for any developer, research team, or organization that wants to evaluate or deploy a 35B-A3B MoE agentic model with open weights — the combination of Apache-2.0 + open weights + 50-author paper + six-domain evaluation + Apple-Silicon-local-friendly .mlx ports + ModelScope mirror makes this a credible alternative to OpenAI / Anthropic / Google proprietary agentic models for teams with the

Who should use it

Any developer or research team that wants to evaluate or deploy an open-weight 35B-A3B MoE agentic model with Apache-2.0 license — the open weights + Apache-2.0 + 50-author paper + six-domain evaluation + Apple-Silicon-local-friendly .mlx ports + ModelScope mirror makes this a credible alternative to OpenAI / Anthropic / Google proprietary agentic models Anyone who values the horizon-scaling argument over the parameter-scaling argument — the paper's core claim is that scaling the agent horizon (long-horizon trajectories + heterogeneous agent abilities via a knowledge-action infrastructure) reaches trillion-parameter-level performance in a 35B-A3B MoE, a meaningful empirical argument against the 'just make it bigger' frontier-model playbook Anyone who needs the multi-teacher domain-routed on-policy distillation recipe — the training recipe is well-documented and reproducible to the degree that any team with 50+ GPUs can replicate it Anyone who needs six-domain evaluation breadth — long-horizon search (HLE, BrowseComp, GAIA), engineering tasks (SWE-Bench-style), scientific research (SciCode, FrontierScience-Olympiad, FrontierScience-Research, HLE, MolBench-bind), instruction following (IFEval, IFBench), general agentic tasks (HiPhO, Seal-0, XBench-DS-2510), scientific agentic tasks (FrontierScience-Research, MolBench-bind) Anyone who needs Apple-Silicon local inference — mlx-community .mlx ports let Apple Silicon users run Agents-A1 on a Mac without an external GPU Anyone who needs Chinese-language developer coverage — the ModelScope mirror serves the Chinese-language developer audience, which is a non-trivial share of the open-weight agentic-modeling community Anyone who needs resource-constrained deployment — Hugging Face quantized variants (GPTQ / AWQ / BNB) cover the practical deployment spectrum: GPTQ/AWQ for NVIDIA, BNB for resource-constrained finetuning Anyone who needs a 35B-A3B MoE with 45K-token average trajectory training data — long enough for internalizing real agentic reasoning, short enough to fit on commodity multi-GPU

Who should skip it

Skip InternScience/Agents-A1 unless the captured evidence suggests it solves a problem you are actively working on.

About this signal

InternScience/Agents-A1 is tracked by RepoRadar as a model release in the New Models section. First seen 2026-07-04; the source record was last checked on 2026-07-04. The current verdict is 'try now' with a Gold tier and moderate setup difficulty. InternScience/Agents-A1 leads on workflow potential (9.2) and practical usefulness (9.0); its lowest signal is setup ease (6.4), so factor that in before investing setup time. This page summarizes the public evidence on the linked source page and states where additional review is still needed.

How this item is evaluated

The InternScience/Agents-A1 record combines a 8.1/10 composite score with separate popularity (1.0), risk (low), and setup (moderate) signals. See the scoring methodology for the current weights and evidence definitions.

Putting this into practice? Read How to vet an AI agent or MCP server before you wire it in for the checklist behind this score.

Risk explanation

README's comparison table (the `Larger-scale Models` column in the per-benchmark tables) references model names that are not currently publicly released as of the cycle date (`GPT-5.5(xhigh)`, `DeepSeek-V4-pro(Max)`, `Kimi-K2.6`); the cycle 146 fictional-model-name-forward-looking rule treats this as a `risk_flag` + `conditional` verdict because the project is a real, runnable, open-weight 35B-A3B MoE agentic model that ships reproducible multi-domain benchmarks against the currently public; Apache-2.0 license with open weights + 50-author paper + Hugging Face collection + ModelScope mirror + mlx-community ports — the combination is the right open-source shape; verify the LICENSE before any commercial embedding that might trigger the cycle 126 'Modified Apache with commercial-use caveat' pattern (this repo is plain Apache 2.0, confirmed 2026-07-04).

Evidence links
Closest alternatives / related signals
open-weight-model agentic-model 35b-moe 35b-a3b moe-mixture-of-experts horizon-scaling long-horizon-trajectories heterogeneous-agent-abilities
Verification record

What RepoRadar actually verified

Discovered

Automated discovery and source capture. Last checked 2026-08-13T20:20:15Z.

No editorial or hands-on review is claimed. This record remains at Discovered.

Verification sources

Longitudinal intelligence

How this decision record is moving

Raw history JSON →

32 dated snapshots retained from 2026-07-04 through 2026-08-13; see the snapshot index for explicit coverage gaps. Stars, version, release, pricing, integration, risk, maintenance, verdict, score, and momentum fields remain explicit even when a source has not reported them. Repository momentum is a normalized 0–10 RepoRadar signal; GitHub stars appear only where the popularity monitor retained exact timestamped observations.

RepoRadar score8.1 current · +0.0 net
Repository momentum7.5 current · -0.5 net
GitHub stars (observed)537 current · +83 net
GitHub stars537 exact observation
VersionNot reported by source
Last releaseNot reported by source
Maintenanceactive
Current risklow
Current verdicttry now
Pricing baselineNo structured commercial pricing baseline
Pricing checkedNot applicable or not recorded
Pricing freshnessNo dated commercial pricing review
Integrations baselineNo structured integrations recorded

Recent dated points

DateScoreMomentumStarsRiskVerdictMaintenance
2026-08-138.17.5537lowtry nowactive
2026-08-128.17.5536lowtry nowactive
2026-08-118.17.8536lowtry nowactive
2026-08-108.17.5533lowtry nowactive
2026-08-098.17.5533lowtry nowactive
2026-08-088.17.8533lowtry nowactive
2026-08-078.17.5528lowtry nowactive
2026-08-068.18.0Not recordedlowtry nownot recorded
2026-08-058.18.0Not recordedlowtry nownot recorded
2026-08-048.17.5528lowtry nowactive
2026-08-038.17.5528lowtry nowactive
2026-08-028.17.5528lowtry nowactive

Why the record changed

stars changed

Stars changed: 536 → 537.

stars changed

Stars changed: 533 → 536.

stars changed

Stars changed: 528 → 533.

stars changed

Source-observed stars changed: 527 → 528. This reports the retained observation delta and does not infer why the upstream change occurred.

stars changed

Stars changed: 475 → 523.

stars changed

Stars changed: 454 → 475.