※重複して受け取られた場合はご容赦下さい
関連研究者の皆様,
NTTの東中です.
NTCIR における対話システムに関するタスクのお知らせです.ご興味ありましたら是非エ
ントリください.多数の参加登録をお待ちしております.
=============================================================
■NTCIR Short Text Conversation (STC) 日本語タスクのお知らせ
(* English version will follow)
この度,NTCIR にて,Short Text Conversation (STC) 日本語タスクを実施することにな
りましたので,参加者募集のご案内をさせていただきます.情報検索,対話システム,自
然言語処理に関わる皆様にぜひご参加いただければと思います.
NTCIR STC URL: http://ntcir12.noahlab.com.hk/stc.htm
■タスク概要
本タスクは,入力ツイートに対して,所定のツイート群から,対話システムの出力として
ふさわしいツイートを抽出するタスクです.
たとえば,「こんにちは」には「こんにちは」と挨拶を返したり,「疲れたよ」には「元
気出して!」「最近お疲れみたいですね」といった相手に寄り添う発言をしたり,「温泉
に行きたいな」には「道後温泉がお勧めですよ」「○○温泉は泉質もよくってもう一度行
きたいです」といった相手の役に立つ発言を抽出するといったことを目指します.広範な
話題を持つツイッターから,対話システムの発話として相応しい発言が抽出できれば,対
話システムの発話能力は飛躍的に高まると考えられます.
なお,日本語タスクでは Twitter のデータを用いますが,中国語タスクでは Weibo を用
いています.中国語タスクの詳細は http://ntcir12.noahlab.com.hk/stc.htm をご覧くだ
さい.
■評価方法
抽出されたツイートについて,人手による主観評価を行います.評価はオーガナイザ側が
行います.主観評価は,0(応答として適合しない), 1(文脈により適合する), 2(適合
する)の3段階のラベルを複数人でラベル付けし,適合するツイートの割合や情報検索の評
価尺度によって評価します.
■データについて
入力対象となるデータは,ランダムにサンプリングされたツイートデータです.2015年の
ツイートデータからランダムサンプリングされたものです.
抽出対象となるデータは,入力対象となるデータとは異なる期間のツイートペア集合(約
100万ツイート)です.これらは,2014年のツイートからランダムに抽出されたものです.
オーガナイザからは,以下のデータを配布します.
・抽出対象ツイートデータ:応答ツイートを抽出する元となるデータ(約100万ツイート)
・開発用データ:ベースライン手法によって抽出されたツイートに対する主観評価値
(各ツイート10名により評価されたものです)
これらは,データ配布サイト https://github.com/mynlp/stc にて公開します.
フォーマルラン時には,別途テストデータを配布します.
ツイートデータは,全てツイートID(id_str)として配布されます.ツイートの本文は参加
者各自にクロールいただくことになりますが,NTTデータ社から一括して購入することも可
能です.
本データを購入されたい方はNTTデータ社 ソーシャルビジネス推進室森様に「NTCIR STC
日本語タスク用データセット」購入のご連絡をお願いします.購入のためのメールアドレ
スは info<at>nazuki-oto.com です.
■参加方法
NTCIR の公式サイトから参加登録をお願いします.
こちらに入力された連絡先アドレスに対して,テストデータ配布等の連絡を行います.
http://ntcir.nii.ac.jp/jp/NTCIR12Regist/
(参加登録は下記〆切まで随時受け付けています)
■スケジュール
検索対象ツイートデータ配布:配布中
開発用ラベル付きデータ配布:11/20
参加登録〆切:2016/1/15
テストデータ配布:2/15
フォーマルラン〆切:2/22
評価結果配布:3/10
論文第一稿〆切:3/20
カメラレディ論文締切:5/1
NTCIR12カンファレンス:6/7-6/10
■お問い合わせ先
日本語タスクについては下記までお願いします.
日本語タスクオーガナイザ
・東中竜一郎(higashinaka.ryuichiro<at>lab.ntt.co.jp)
・宮尾祐介(yusuke<at>nii.ac.jp)
NTCIR STC 全体にかかわるお問い合わせについては以下までお願いします.
・STCメーリングリスト(ntcirstc-organizer<at>yahoogroups.com)
本メーリングリストにはオーガナイザ全員(以下)が含まれます.
Hang Li, Noah's Ark Lab, Huawei, Hong Kong
Tetsuya Sakai, Waseda University, Japan
Zhengdong Lu, Noah's Ark Lab, Huawei, Hong Kong
Lifeng Shang, Noah's Ark Lab, Huawei, Hong Kong
Yusuke Miyao, National Institute of Informatics, Japan
Ryuichiro Higashinaka, Nippon Telegraph and Telephone Corporation, Japan
=============================================================
Call for participation: NTCIR Short Text Conversation (STC) Japanese task
This is a call for participation for the NTCIR Short Text Conversation (STC)
Japanese task. We ask those who are working in the field of information
retrieval, dialogue systems, and natural language processing in general to
participate in the task.
*Task description
Given an input tweet, the task is to retrieve, from a pool of tweets, tweets
that are suitable as responses for a dialogue system. Since Twitter contains
utterances on a wide variety of topics, if we can extract reasonable tweets as
responses, it will lead to the improvement in the language generation capability
of dialogue systems.
Note that, in the Japanese task, Twitter is used; however, in the Chinese task,
Weibo is used. See http://ntcir12.noahlab.com.hk/stc.htm for the details of the
Chinese task.
*Evaluation
The extracted tweets will be evaluated by human judges. The evaluation is done
by the organizers. In the evaluation process, each extracted tweet is labeled
with 0 (inappropriate), 1 (appropriate in some context), and 2 (appropriate) by
multiple judges. The rate of appropriate answers and information retrieval (IR)
related measures (graded relevance IR measures) will be used as evaluation
metrics.
*Data
The input data will be those randomly sampled from tweets in the year 2015. The
pool of tweets (the target for extraction) is the randomly sampled tweet pairs
(mention-reply pairs) in the year 2014. The size of the pool is just over one
million; that is 500K pairs.
The following data will be provided from the organizers:
(1) Twitter data (by using their IDs) 1M in size
(2) Development data. Input samples and output samples annotated with reference
labels. Here, the number of annotators is ten.
The data can be downloaded from https://github.com/mynlp/stc and the test data
will be provided at the time of the formal run.
Since the Twitter data are provided in the form of IDs, it will be necessary for
the participants to crawl the data by themselves. Alternatively, the
participants can purchase the data from NTT DATA Corporation. For those who
want to purchase the data, please send an email to info<at>nazuki-oto.com,
requesting the data for the NTCIR STC Japanese task.
* How to participate
Please register from the NTCIR official website:
http://ntcir.nii.ac.jp/jp/NTCIR12Regist/
We will inform the participants of any update by using the registered contact
addresses.
* Schedule
- Release of the Twitter data: done
- Release of the development data: 11/20
- Registration deadline: 2016/1/15
- Release of the test data: 2/15
- Formal run deadline: 2/22
- Distribution of evaluation results: 3/10
- Paper draft deadline: 3/20
- (brief review of the draft papers)
- Camera ready deadline: 5/1
- NTCIR12 conference: 6/7-6/10
* Contact
Regarding the Japanese task, please contact the organizers of the Japanese task.
- Ryuichiro Higashinaka (higashinaka.ryuichiro<at>lab.ntt.co.jp)
- Yusuke Miyao (yusuke<at>nii.ac.jp)
For general inquiries related to NTCIR STC, please send an email to
- STC mailing list (ntcirstc-organizer<at>yahoogroups.com)
The mailing list includes all the organizers:
Hang Li, Noah's Ark Lab, Huawei, Hong Kong
Tetsuya Sakai, Waseda University, Japan
Zhengdong Lu, Noah's Ark Lab, Huawei, Hong Kong
Lifeng Shang, Noah's Ark Lab, Huawei, Hong Kong
Yusuke Miyao, National Institute of Informatics, Japan
Ryuichiro Higashinaka, Nippon Telegraph and Telephone Corporation, Japan
--
Ryuichiro Higashinaka
NTT Media Intelligence Laboratories. NTT Corp.
1-1 Hikarinooka, Yokosuka, 239-0847 Japan.
phone: +81-46-859-2027 fax: +81-46-855-1054
日本データベース学会の皆様,
(重複して受け取られた場合にはご容赦ください)
京都大学の加藤と申します.
NTCIR-12 MobileClick-2タスクの参加募集をお送りさせていただきます.
これまでのところ,21チーム,38ユーザにご参加いただき,75のランが提出され
ています.参加登録締切は【2016年2月1日】となっております.
(The English message follows the Japanese one)
______________________________________________________________________
NTCIR-12 MobileClick-2 Task
-- モバイル情報アクセスタスク --
参加チーム募集
http://www.mobileclick.org/?locale=ja (日本語版公開)
評価システムが現在稼働中で2016年2月4日まで提出を受け付けております
参加登録締切: 2016年2月1日
______________________________________________________________________
情報アクセス技術評価ワークショップNTCIR-12のコアタスクの1つである
MobileClick-2では,与えられたクエリに対して簡潔な2層の要約を自動生成する
ことによって,
直接的かつ即時のモバイル情報アクセスを実現することを目的としています.
この要約の1層目は重要かつ概要的な情報を提示し,1層目からのリンクをクリッ
クすることで
閲覧できる2層目は,1層目より詳細な情報を含むことが期待されます.
このMobileClickの目標は2つのサブタスク,iUnit retrievalとiUnit
summarizationサブタスクを通して達成可能であると考えられます.
iUnit rankingサブタスクは,クエリとiUnitと呼ばれる情報の断片が与えられた
時にiUnitをクエリに対する重要度順にランキングするであり,
iUnit summarizationサブタスクは,クエリとiUnit集合,クエリの意図が与えら
れた時にiUnitをうまく配置して2層の要約を生成するタスクとなっています.
===================================
MobileClick-2に参加すべき3つの理由
===================================
* 現代の検索エンジンにとってより重要性を増すモバイル情報アクセス技術を評
価できる
多くの商用検索エンジンが与えられたクエリに対してダイレクトアンサーを提供
し始めており,
それを生成する技術はモバイル検索において非常に重要なチャレンジであると考
えられます.
MobileClick-2ではテストコレクションを提供し,重要な情報の断片を文書から
抽出し,
それらを要約するアルゴリズムを評価することを可能にします.
* 要約,情報抽出,意図推定,検索結果多様化を評価するための大規模テストコ
レクション
これまでのNTCIRタスクにおいて,我々は複数のテストコレクションを構築して
きました.
参加者となればこれらのテストコレクションをクエリ指向要約やスニペット生成
などの評価に利用
することができます.これらに加えて,今回は各クエリに対する意図も提供する
ため,
意図推定や検索結果多様化の評価にも活用いただけます.
また,ヤフー株式会社より本タスクのクエリに関する「検索関連クエリデータ」
が提供されます.
大変貴重なデータですので,ぜひとも参加の上,お役立てください.
* リアルタイム評価システム
今回のタスクではシステム開発を促進し評価結果を共有するためのリアルタイム
評価システムを
導入しております.そのため,ランを提出後にすぐに結果を確認可能であり,ま
た,システムの
有効性を広く周知することができます.まずは,
公式Webサイト http://www.mobileclick.org/?locale=ja
にてアカウントを作成しランを提出してみてください.
===============================
参加登録
===============================
本タスクに興味を持たれた研究グループはぜひNTCIR参加登録ページ
http://ntcir.nii.ac.jp/NTCIR12Regist/ にて参加登録をお願いいたします.
また,公式Webサイト http://www.mobileclick.org/?locale=ja
にてユーザアカウントを作成の上,試しにランを提出してみてください.
===============================
重要な日程
===============================
2016年2月1日 ドラフト概要論文公開
2016年2月4日 ラン投稿締切
2016年3月1日 ドラフト参加者論文提出締切
2016年6月 NTCIR-12 conference
===============================
お問い合わせ
===============================
Mail: ntcadm-1click(a)nii.ac.jp
Twitterアカウントもフォローください: @mobileclicktask
MobileClick-2運営者
(English version)
______________________________________________________________________
NTCIR-12 MobileClick-2 Task
-- Mobile Information Access Task --
Call for Participation
http://www.mobileclick.org/
38 users from 21 teams have registered online and submitted 75 runs.
The evaluation system is now running,
and waiting for your submission by February 4, 2016.
Registration due: February 1, 2016.
______________________________________________________________________
MobileClick-2, a core task at NTCIR-12, aims to provide direct and
immediate mobile information access, by automatically generating
a concise two-layered summary for a given query.
The first layer is expected to contain the most important information
and an outline of additional relevant information, while the second
layer is expected to contain detailed information that can be
accessed by clicking on an associated anchor text in the first layer.
The goal of MobileClick can be achieved through two subtasks:
iUnit ranking subtask and iUnit summarization subtask.
iUnit ranking subtask is a task where systems are expected to rank
a set of pieces of information (iUnits) based on their importance for
a given query.
iUnit summarization subtask is defined as follows: Given a query,
a set of iUnits, and a set of intents, generate a structured textual
output.
===============================
Why should you participate?
===============================
* Evaluate mobile information access technologies that have been
becoming more important in modern search engines
Direct answers in response to a given query are provided by many
popular commercial search engines, and are considered as important
challenges to achieve immediate mobile information access.
MobileClick-2 offers test collections, by which you can evaluate
your algorithms that identify important information pieces and summarize
them for a given search query.
* Large-scale test collections are available for evaluating summarization,
information extraction, intent detection, and search result diversification
We have developed several test collections in a series of NTCIR tasks.
Participants can utilize these test collections to evaluate query-focused
summarization, search result snippet generation, etc. In addition, we also
provide intents of each query, which can be used to evaluate intent
detection and result diversification techniques.
* Real time evaluation system
We provide a real time evaluation system that accelarates your system
development and shares records among participants. You can immediately
get results and publicly demonstrate the effectiveness of your system.
Create your account and submit your runs at http://www.mobileclick.org/ .
===============================
Registration
===============================
Participation is open to all interested research groups.
Please visit http://ntcir.nii.ac.jp/NTCIR12Regist/ for registration.
You may also create a user account at http://www.mobileclick.org/
to obtain test collections and submit your runs.
===============================
Schedule
===============================
Feb 1, 2016 Early draft task overview release
Feb 4, 2016 Run submission due
Mar 1, 2016 Draft participant paper submission due
Jun 2016 NTCIR-12 conference
===============================
Contact
===============================
Mail: ntcadm-1click(a)nii.ac.jp
Please follow us on twitter @mobileclicktask
Best regards,
Makoto P. Kato
on behalf of the MobileClick-2 organizers