Open-Vocabulary Object Detection(OVOD)는 학습 시 사용된 카테고리에만 한정되는 기존 객체 탐지 방식의 한계를 극복하기 위해 제안된 기법이다. 기존 OVOD는 탐지하고자 하는 물체를 “a {category}” 라...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
https://www.riss.kr/link?id=A109751637
2025
Korean
KCI우수등재
학술저널
499-507(9쪽)
0
상세조회0
다운로드Open-Vocabulary Object Detection(OVOD)는 학습 시 사용된 카테고리에만 한정되는 기존 객체 탐지 방식의 한계를 극복하기 위해 제안된 기법이다. 기존 OVOD는 탐지하고자 하는 물체를 “a {category}” 라...
Open-Vocabulary Object Detection(OVOD)는 학습 시 사용된 카테고리에만 한정되는 기존 객체 탐지 방식의 한계를 극복하기 위해 제안된 기법이다. 기존 OVOD는 탐지하고자 하는 물체를 “a {category}” 라는 프롬프트를 활용해 분류기를 생성하여 물체를 탐지하였으나, 본 논문에서는 탐지하고자 하는 물체의 계층적 구조를 프롬프트에 적용하여 탐지 능력을 향상하였다. 특히, 문장의 길이가 길어지는 연결어의 사용을 줄이고, 강조하고자 하는 단어를 문장 앞에 위치시키는 등의 프롬프트 엔지니어링 방식을 사용하여 더 좋은 탐지 성능을 가지는 것을 확인하였다. 이는 물체의 계층 구조에 따른 내재적 의미를 잘 나타내는 문장을 구성할 수 있으며, 추가적인 컴퓨팅 자원 없이 분류기를 생성할 수 있다는 장점을 지닌다. 또한, 이미지 캡셔닝, 의료 영상 분석 등의 분야에서도 적용 가능하며, 사람에게 익숙한 계층적 표현을 활용함으로써 모델의 설명력 향상에 기여할 수 있다.
다국어 초록 (Multilingual Abstract)
Open-Vocabulary Object Detection (OVOD) has been proposed to overcome the limitation of traditional object detection methods, which are restricted to recognizing only categories seen during training. While conventional OVOD approaches generate classif...
Open-Vocabulary Object Detection (OVOD) has been proposed to overcome the limitation of traditional object detection methods, which are restricted to recognizing only categories seen during training. While conventional OVOD approaches generate classifiers using simple prompts like “a {category}”, this paper incorporates the hierarchical structure of object categories into prompts to enhances detection performance. Specifically, we applied prompt engineering techniques that could reduce the use of lengthy connectives and place important keywords at the beginning of the sentence. This resulted in more effective prompts that could capture the intrinsic meaning of hierarchical information. Our method allows for the generation of classifiers without additional computational resources or retraining. Furthermore, it demonstrates strong generalizability. It can be applied to other tasks such as image captioning and medical image analysis. By leveraging hierarchical expressions familiar to humans, our approach also contributes to improving the interpretability of model outputs.
전기식 RTO 에너지 소모 최소화를 위한 PFD 시뮬레이터 기반 심층 강화학습
대규모 언어 모델에 기반한 질문 재작성을 활용한 검색증강 생성 시스템
텍스트 부가 정보를 활용한 선형 기반 순차적 추천 모델