정보관리학회지, 한국정보관리학회

21

유재복(한국원자력연구원) ; 정영미(연세대학교) 2010, Vol.27, No.1, pp.103-118 https://doi.org/10.3743/KOSIM.2010.27.1.103

초록보기

초록

최근 특허기술의 가치평가가 크게 강조되고 있으며, 그 평가의 수단으로 특허의 피인용횟수가 매우 유용한 척도 중의 하나로 받아들여지고 있다. 그에 따라 이 연구에서는 특허의 피인용횟수와 이에 영향을 미칠만한 형태적․기술적․개념적 요인의 17개 변수들 간의 상관관계를 미국특허를 대상으로 5개 주제분야에 걸쳐 분석하였다. 분석결과 특허의 피인용횟수와 일정 수준 이상의 상관관계, 즉 5% 이상의 설명력을 갖는 변수는 페이지 수, 청구항 수, 참고문헌 평균 피인용횟수, 기술분야 특허증감율, 서지결합도, 동시인용도 및 문헌간유사도 등 7개로 나타났다. 또한 이들 변수에 대한 분산분석 결과 7개 변수 모두 전반적으로 대부분의 주제분야 간에 있어서 평균값의 차이가 있는 것으로 나타났다.

Abstract

Recently, the valuation of patented technology has been greatly emphasized, and patent citation has been accepted as a very useful index of this technology. In this study, we performed correlation analyses between the patent citation counts and 17 explanatory variables of morphological, technological, and conceptual factors with a test dataset of U.S. patents in five subject fields. Seven variables having 5% or more standardized variances(r2) with patent citation counts were identified; number of pages, number of claims, reference-average-citation rate, patent increase/decrease rate, strength of bibliographic coupling, co-citation counts and document similarity. The result of the ANOVA test shows that the mean values of these variables vary among most subject fields.

22

사건중심 뉴스기사 자동요약을 위한 사건탐지 기법에 관한 연구

정영미(연세대학교) ; 김용광(연세대학교) 2008, Vol.25, No.4, pp.227-243 https://doi.org/10.3743/KOSIM.2008.25.4.227

초록보기

초록

이 연구에서는 사건중심 뉴스기사 요약문을 자동생성하기 위해 뉴스기사들을 SVM 분류기를 이용하여 사건 주제범주로 먼저 분류한 후, 각 주제범주 내에서 싱글패스 클러스터링 알고리즘을 통해 특정한 사건 관련 기사들을 탐지하는 기법을 제안하였다. 사건탐지 성능을 높이기 위해 고유명사에 가중치를 부여하고, 뉴스의 발생시간을 고려한 시간벌점함수를 제안하였다. 또한 일정 규모 이상의 클러스터를 분할하여 적절한 크기의 사건 클러스터를 생성하도록 수정된 싱글패스 알고리즘을 사용하였다. 이 연구에서 제안한 사건탐지 기법의 성능은 단순 싱글패스 클러스터링 기법에 비해 정확률, 재현율, F-척도에서 각각 37.1%, 0.1%, 35.4%의 성능 향상률을 보였고, 오보율과 탐지비용에서는 각각 74.7%, 11.3%의 향상률을 나타냈다.

Abstract

This study investigates an event detection method with the aim of generating an event-focused news summary from a set of news articles on a certain event using a multi-document summarization technique. The event detection method first classifies news articles into the event related topic categories by employing a SVM classifier and then creates event clusters containing news articles on an event by a modified single pass clustering algorithm. The clustering algorithm applies a time penalty function as well as cluster partitioning to enhance the clustering performance. It was found that the event detection method proposed in this study showed a satisfactory performance in terms of both the F-measure and the detection cost.

23

과학기술분야 국제협력 증진을 위한 아시아 국가 간 공동연구 현황 분석

김원진(연세대학교) ; 정영미(연세대학교) 2010, Vol.27, No.3, pp.103-123 https://doi.org/10.3743/KOSIM.2010.27.3.103

초록보기

초록

과학기술분야 국제협력은 국가 경쟁력 확보를 위해서 필수적이다. 한국은 과학기술의 인적․물적 자원의 한계를 극복하고자 연구의 국제화를 추진하고 있으며 최근 아시아 국가와 연구협력에서 높은 성장률을 보여주었다. 본 연구에서는 네트워크 분석을 이용하여 한국과의 공동연구가 크게 증가한 아시아 국가 간 공동연구 현황을 공저논문 수와 주제범주로 구분하여 실증적으로 파악하였다. 최근 5년간 아시아 국가 간 공저논문 수 기반 네트워크를 살펴보면, 일본, 중국, 한국 등 동북아시아 국가들이 네트워크 중심부에 있었으며 국가 상호 간 공동연구가 활발하게 이루어졌다. 또한 아시아 지역별로 공동연구의 주제범주를 분석한 결과, 동북아시아 지역은 기초과학 분야에서, 남부아시아, 동남아시아, 서남아시아 지역은 의학 분야에서 공동연구가 활발하게 이루어진 것으로 나타났다.

Abstract

Recently, research community in Korea has shown a rapid growth in collaborating with Asian countries. In this study, we analyzed research collaboration among Asian countries using network analysis of co-authored papers as well as subject categories. The network of co-authored papers among Asian countries over the 5-year period since 2005 revealed that Japan, China, and Korea were positioned at the central part of the network and highly productive in collaborative research. In the analysis of the subject categories of co-authored papers in four different Asian regions with 2009 data, physics and material science were found the most productive subject fields in collaborative research in Northeast Asia. On the other hand, medical science was the most collaborative subject field in the remaining Asian regions.

24

지구적 환경문제 해결을 위한 학술활동과 환경운동 경향 연구

박재신(연세대학교) ; 정영미(연세대학교) 2010, Vol.27, No.3, pp.83-102 https://doi.org/10.3743/KOSIM.2010.27.3.083

초록보기

초록

본 연구에서는 지구적 환경문제의 해결 방식으로서 환경과학 분야의 학술활동과 같은 학문적 접근 방식과 환경 NGO 중심의 환경운동과 같은 실천적 접근 방식을 두 가지 주요 흐름이라 보고, 이들 각각의 특성을 계량정보학적 분석을 통해 파악하고 비교하였다. 지난 10년 간 환경과학 분야에서 인용된 저널의 주제범주 간 동시인용 관계를 분석함으로써 이 분야의 지식 구조를 파악하였고, 환경 NGO의 웹 사이트에서 수집된 외부링크 데이터를 이용하여 이들의 관심 분야를 확인하였다. 또한 저널 논문과 NGO 뉴스에서 추출된 핵심어를 이용한 동시출현단어 분석을 통해 하위 주제를 파악하여 이들 간의 주제적 유사성과 상이성을 구체화하였다.

Abstract

This study aims to understand and compare the characteristics of two major approaches to solving global environmental problems-an academic approach including scholarly activities of environmental sciences and a practical approach of environmental movements led by NGOs-by employing informetric analysis methods. Knowledge structure of environmental sciences is depicted through co-citation networks of subject categories assigned to the cited journals in the discipline of environmental sciences for the 10-year period from 2000 to 2009. Furthermore, major interests of environmental NGOs are identified on the basis of external link data collected from web sites of the NGOs. Co-word analyses are also performed using the texts of journal papers in environmental sciences as well as news articles provided by NGO sites. Through the analyses, dominant subject areas of environmental sciences and environmental movements are identified demonstrating similarities and differences between the two approaches.

25

특허인용 예측모형 구축에 관한 연구

유재복(한국원자력연구원) ; 정영미(연세대학교) 2010, Vol.27, No.4, pp.239-258 https://doi.org/10.3743/KOSIM.2010.27.4.239

초록보기

초록

이 연구에서는 특허의 인용에 영향을 미치는 주요 변수들을 토대로 특허의 피인용횟수를 예측하기 위한 모형을 제시하였다. 이를 위해 미국특허를 대상으로 5개 주제분야에 걸쳐 특허의 피인용횟수와 일정 수준 이상의 상관관계, 즉 5% 이상의 설명력을 갖는 것으로 밝혀진 페이지 수, 청구항 수, 참고문헌 평균 피인용횟수, 서지결합도, 문헌간유사도 등 5개 변수들을 토대로 다중회귀분석을 실시하였다. 연구결과에 따르면, 제시된 5개 주제분야의 특허인용 예측모형의 설명력은 주제분야에 따라 58.3%~89.6%로 나타났으며, 예측변수로 사용된 5개의 독립변수 중 특허 피인용횟수에 가장 영향력이 높은 변수는 ‘문헌간유사도’로 나타났다. 또한 이 연구에서 추정된 주제분야별 예측모형을 토대로 산출한 특허 피인용횟수에 대한 예측값과 실제값을 비교한 결과 이들 예측모형은 5개 주제분야에서 모두 적합한 것으로 나타났다.

Abstract

The purpose of this study is to develop a prediction model of patent citation counts based on major factors which affect patent citation. To this end, we performed multiple regression analysis between the patent citation counts and five explanatory variables such as the number of pages, the number of claims, the reference-average-citation rate, the strength of bibliographic coupling, and the document similarity proved as having 5% or more standardized variances(r2) with patent citation counts, with a test dataset of U.S. patents in five subject fields. As a result, our prediction models showed 58.3% to 89.6% predictability depending on subject fields and revealed the document similarity has the highest impact on citation counts among the five predictive variables in all the subject fields. The result of comparison between the predicted citation counts and the actual ones confirmed the usefulness of the citation prediction models built for each subject field.

26

학습문헌집합에 기 부여된 범주의 정확성과 문헌 범주화 성능

심경(Systems R&D Center, Iris.Net) ; 정영미(연세대학교) 2006, Vol.23, No.2, pp.265-285 https://doi.org/10.3743/KOSIM.2006.23.2.265

초록보기

초록

문헌범주화에서는 학습문헌집합에 부여된 주제범주의 정확성이 일정 수준을 가진다고 가정한다. 그러나, 이는 실제 문헌집단에 대한 지식이 없이 이루어진 가정이다. 본 연구는 실제 문헌집단에서 기 부여된 주제범주의 정확성의 수준을 알아보고, 학습문헌집합에 기 부여된 주제범주의 정확도와 문헌범주화 성능과의 관계를 확인하려고 시도하였다. 특히, 학습문헌집합에 부여된 주제범주의 질을 수작업 재색인을 통하여 향상시킴으로써 어느 정도까지 범주화 성능을 향상시킬 수 있는가를 파악하고자 하였다. 이를 위하여 과학기술분야의 1,150 초록 레코드 1,150건을 전문가 집단을 활용하여 재색인한 후, 15개의 중복문헌을 제거하고 907개의 학습문헌집합과 227개의 실험문헌집합으로 나누었다. 이들을 초기문헌집단, Recat-1, Recat-2의 재 색인 이전과 이후 문헌집단의 범주화 성능을 kNN 분류기를 이용하여 비교하였다. 초기문헌집단의 범주부여 평균 정확성은 16%였으며, 이 문헌집단의 범주화 성능은 F1값으로 17%였다. 반면, 주제범주의 정확성을 향상시킨 Recat-1 집단은 F1값 61%로 초기문헌집단의 성능을 3.6배나 향상시켰다.

Abstract

In text categorization a certain level of correctness of labels assigned to training documents is assumed without solid knowledge on that of real-world collections. Our research attempts to explore the quality of pre-assigned subject categories in a real-world collection, and to identify the relationship between the quality of category assignment in training set and text categorization performance. Particularly, we are interested in to what extent the performance can be improved by enhancing the quality (i.e., correctness) of category assignment in training documents. A collection of 1,150 abstracts in computer science is re-classified by an expert group, and divided into 907 training documents and 227 test documents (15 duplicates are removed). The performances of before and after re-classification groups, called Initial set and Recat-1/Recat-2 sets respectively, are compared using a kNN classifier. The average correctness of subject categories in the Initial set is 16%, and the categorization performance with the Initial set shows 17% in F1 value. On the other hand, the Recat-1 set scores F1 value of 61%, which is 3.6 times higher than that of the Initial set.

27

온라인 자료 식별체계 실태조사를 기반으로 한 납본연계방안 제안 연구

노영희(건국대학교 문헌정보학과 교수) ; 손애경(글로벌사이버대학교 미디어콘텐츠창작학과 교수) ; 이경선(서강대학교 공공정책대학원 행정법무학과 교수) ; 장인호(대진대학교 문헌정보학과 부교수) ; 정영미(동의대학교 문헌정보학과 교수) ; 차현주(성균관대학교 문헌정보학과 초빙교수) 2024, Vol.41, No.1, pp.133-162 https://doi.org/10.3743/KOSIM.2024.41.1.133

초록보기

초록

디지털화가 급속히 진행됨에 따라, 온라인 자료의 식별 및 관리의 중요성이 대두되고 있다. 특히, 디지털 콘텐츠의 효율적인 유통 및 보존을 위한 체계적인 식별체계의 필요성이 증가하고 있다. 본 연구는 이러한 시대적 요구에 부응하여, 온라인 자료의 식별 및 관리를 위한 현행 식별체계의 실태를 조사하고, 이를 납본과 연계하여 보다 체계적인 관리 및 활용 방안을 모색하는 것을 목적으로 한다. 이를 위해 온라인 자료 식별체계와 발급실태를 조사하고 온라인 자료에 관련된 선행연구를 분석하였다. 분석결과를 기반으로 한 납본 연계방안은 다음과 같이 세 가지로 요약할 수 있다. 첫째, 납본의 우선순위 및 활용성을 위해 납본과 이용의 상호보완 강화, 납본의 우선순위 부여, 납본자료의 활용성 증대 전략이 요구된다. 둘째, 국제표준번호를 기반으로 한 납본 연계 방안으로서, ISBN과 UCI의 연계 납본, 국제표준자료번호와 납본 연계, 국제표준번호와 UCI의 메타데이터연계, UCI와 ICN의 연계 통합, 납본시스템 고도화를 위한 자동화 기술 도입 전략이 요구된다. 셋째, 위에서 제안한 전략들이 그 효과적으로 작용하기 위해서는 정책적인 지원도 같이 이루어져야 할 것이다. 한국서지표준센터의 납본 역할 강화를 포함하여 출판사와의 협력강화, 납본자료에 대한 보상, 납본제도에 대한 인식 제고 및 제도적 보상 등의 측면에서 고려되어야 할 부분이 있다.

Abstract

The rapid digitalization has highlighted the importance of identifying and managing online resources. Especially, the need for a systematic identification system for the efficient distribution and preservation of digital content is growing. This study aims to respond to these contemporary demands by investigating the current state of identification systems for online resources and exploring more systematic management and utilization methods through linking these systems with legal deposit. To achieve this, the study surveyed the identification systems and their issuance status for online resources and analyzed prior research related to these online resources. Based on the analysis, the proposed strategies for linking with legal deposit can be summarized into three categories: First, to prioritize and enhance the utilization of legal deposit, strategies are required to strengthen the mutual complementarity of deposit and use, to assign priorities to certain deposits, and to increase the usability of deposited materials. Second, as strategies based on international standard numbers for linking with legal deposit, it is necessary to integrate ISBN and UCI in the deposit process, to link international standard resource numbers with deposit, to interconnect metadata between international standard numbers and UCI, to integrate UCI and ICN, and to introduce automation technology for upgrading the deposit system. Third, to effectively implement the aforementioned strategies, policy support is essential. This includes enhancing the role of the Korean Bibliographic Standards Center, strengthening cooperation with publishers, compensating for deposited materials, and increasing awareness and institutional compensation for the legal deposit system.

바로가기메뉴

초록

Abstract

초록

Abstract

초록

Abstract

초록

Abstract

초록

Abstract

초록

Abstract

초록

Abstract

정보관리학회지