SAVRN
Search Contact SAVRN

Open-weight model · Object detection

ORena-SurgHint-solution

by Orhun Utku Aydin OUAydin/ORena-SurgHint-solution

SurgHint is a surgical visual question answering system developed for the FRAME and SEGMENT tracks of the ORena FOCUS Challenge.

Parameters
Context
Weights1.3 GB
Licenseother
AccessAccess requested at publisher
Monthly Downloads

Model Card

SurgHint is a surgical visual question answering system developed for the FRAME and SEGMENT tracks of the ORena FOCUS Challenge. This repository provides the detector checkpoints and Qwen LoRA adapters for answering questions about foreign objects in surgical images and video clips. VLMMAXXING — Orhun Utku Aydin, Frank te Nijenhuis, Dietmar Frey. Sign in and request access on this model page to download the checkpoints. Qwen base weights and FRAME's frozen DINOv3 backbone are not included; please obtain them separately under their upstream licences and access terms. SurgHint inference uses greedy decoding (dosample=False) with thinking disabled. The inference code overrides the sampling…

Excerpt from the card by Orhun Utku Aydin, licensed other.

Identity and Version

Repository
OUAydin/ORena-SurgHint-solution
Publisher
Orhun Utku Aydin
Task
Object detection
Modality
Image
Library
Not stated by the source
Parameters
Not stated by the source
Languages
en
Revision
e1430fc4e8e5b5653eb13d91d98afb74e2c654bb
First published
2026-09-18
Last updated
2026-09-18

Files and Weights

27 files, 1.3 GB in total. The weights are 4 files totalling 1.3 GB in pth, safetensors.

Weights4 files · 1.3 GB
Configuration5 files · 7.7 KB
Tokenizer2 files · 20.0 MB
Documentation8 files · 42.2 KB
Other6 files · 54.5 KB
Repository2 files · 1.6 KB
Every file
FileTypeSizeSHA-256
FRAME-track/weights/custom/plain-detr-head.pthWeights514.6 MB
FRAME-track/weights/custom/qwen-lora/adapter_model.safetensorsWeights205.2 MB
SEGMENT-track/weights/custom/qwen-lora/adapter_model.safetensorsWeights467.1 MB
SEGMENT-track/weights/custom/rfdetr-large.pthWeights136.6 MB
FRAME-track/weights/custom/qwen-lora/adapter_config.jsonConfiguration1.3 KB
SEGMENT-track/weights/custom/qwen-lora/adapter_config.jsonConfiguration1.2 KB
SEGMENT-track/weights/custom/qwen-lora/config.jsonConfiguration3.8 KB
SEGMENT-track/weights/custom/qwen-lora/generation_config.jsonConfiguration200 B
SEGMENT-track/weights/custom/qwen-lora/processor_config.jsonConfiguration1.2 KB
FRAME-track/weights/custom/DINOv3-LICENSE.mdDocumentation7.5 KB
FRAME-track/weights/custom/qwen-lora/LICENSE-APACHE-2.0.txtDocumentation11.5 KB
FRAME-track/weights/custom/qwen-lora/NOTICE.mdDocumentation698 B
LICENSES/DINOv3-LICENSE.mdDocumentation7.5 KB
README.mdDocumentation2.3 KB
SEGMENT-track/weights/custom/qwen-lora/LICENSE-APACHE-2.0.txtDocumentation11.3 KB
SEGMENT-track/weights/custom/qwen-lora/NOTICE.mdDocumentation864 B
SEGMENT-track/weights/custom/rfdetr-large.pth.NOTICE.mdDocumentation459 B
LICENSES/Apache-2.0-Qwen3.5-9B.txtOther11.5 KB
LICENSES/Apache-2.0-Qwen3.6-27B.txtOther11.3 KB
LICENSES/Apache-2.0-RF-DETR.txtOther11.3 KB
LICENSES/MIT-Plain-DETR.txtOther1.1 KB
SEGMENT-track/weights/custom/Apache-2.0-RF-DETR.txtOther11.3 KB
SEGMENT-track/weights/custom/qwen-lora/chat_template.jinjaOther7.8 KB
.gitattributesRepository1.6 KB
.gitignoreRepository22 B
SEGMENT-track/weights/custom/qwen-lora/tokenizer.jsonTokenizer20.0 MB
SEGMENT-track/weights/custom/qwen-lora/tokenizer_config.jsonTokenizer1.2 KB

License and Download

License
other
Access
Access requested at publisher
Download size
1.3 GB
Request access from Orhun Utku Aydin

Orhun Utku Aydin grants access through its official repository on Hugging Face. Read the license.

Memory Requirements

PrecisionWeights in memory
As published1.3 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About ORena-SurgHint-solution

What license is ORena-SurgHint-solution released under?

other, as its publisher declares it. Read the license text before commercial use.

Similar Models

Model · Object detection

locate-anything.cpp-gguf

Mudler

GGUF builds of nvidia/LocateAnything-3B for locate-anything.cpp - a C++/ggml inference engine for open-vocabulary detection / visual grounding, no Python at inference time. Brought to you by the LocalAI team. The detections are the same as the official PyTorch implementation (the engine is parity-gated against it), and it runs faster - on CPU and GPU. The full-precision f32 GGUF (~15 GB) is reproducible from the HF weights with scripts/convertlocateanythingtogguf.py in the repo. Same detections as the official model, faster. Full methodology, the warm/median setup, parity checks, and more images are in the repo's Slow-mode inference on the 448 fixture; vs official divides the official…

Open weights other gguf

Model · Object detection

Anzhcs_YOLOs

Anzhc

YOLOs in this repo are trained with datasets that i have annotated myself, or with the help of my friends(They will be appropriately mentioned in those cases). YOLOs on open datasets will have their own pages. Ultralytics 8.3.217 updates mask handling, which breaks function in main Adetailer repo. Install Ultralytics==8.3.216 or lower. Alternatively - use forks that fix this. - Fixed in main repo. I've added some features to make Adetailer more usable and less manual - https://github.com/Anzhc/aadetailer-reforge Im open to commissions, hit me up in Discord - anzhc P.S. All model names in tables have download links attached:3 Series of models aiming at detecting and segmenting face…

Open weights agpl-3.0 ultralytics

This is a fine-tuned version of YOLOv11 (n, s, m, l, x) specialized for License Plate Detection, using a public dataset from Roboflow Universe: The upstream Roboflow dataset (license-plate-recognition-rxg4e) contains train/test contamination — the same source images appear in both the training and test splits with only minor manual augmentation applied (see Discussion #2 for concrete examples). As a result: - The reported metrics below are likely overestimated, because the test set is not a true held-out evaluation. - Real-world generalization performance is expected to be lower than the numbers in the table. - Treat all evaluation figures with caution and validate the model on your own…

Open weights agpl-3.0 ultralytics

Model · Object detection

surya_layout2

Datalab

A lightweight document layout detection model used by Surya. It detects layout regions (text, tables, figures, headers, captions, equations, etc.) on a page image and runs on CPU or GPU. This is the "fast" layout detector — a compact object detector that serves as a drop-in alternative to Surya's VLM-based layout model. Documentation, installation, and everything else lives in the Point the fast layout predictor at this checkpoint: Or make it the default so the CLI and library use it without an explicit path: Released under the AI Pubs OpenRAIL-M license (see LICENSE) — the same license as the surya-ocr-2 model weights.

Open weights openrail surya

Model · Object detection

detr-resnet-50

Joshua

https://huggingface.co/facebook/detr-resnet-50 with ONNX weights to be compatible with Transformers.js. If you haven't already, you can install the Transformers.js JavaScript library from NPM using: Test it out here, or create your own object-detection demo with 1 click! Note: Having a separate repo for ONNX weights is intended to be a temporary solution until WebML gains more traction. If you would like to make your models web-ready, we recommend converting to ONNX using Optimum and structuring your repo like this one (with ONNX weights located in a subfolder named onnx).

Open weights 1,024 tokens transformers.js