Lucene-2.2.0 源代码阅读学习(27)

xiaoxiao2026-09-14  15

关于Lucene的检索(IndexSearcher)的内容。

通过一个例子,然后从例子所涉及到的内容出发,一点点地仔细研究每个类的实现和用法。

先写一个简单的使用Lucene实现的能够检索的类,如下所示:

package org.shirdrn.lucene;

import java.io.IOException;import java.util.Date;import java.util.List;

import org.apache.lucene.document.Document;import org.apache.lucene.document.Fieldable;import org.apache.lucene.index.CorruptIndexException;import org.apache.lucene.index.IndexReader;import org.apache.lucene.index.Term;import org.apache.lucene.index.TermDocs;import org.apache.lucene.search.Filter;import org.apache.lucene.search.Hits;import org.apache.lucene.search.IndexSearcher;import org.apache.lucene.search.Query;import org.apache.lucene.search.TermQuery;import org.apache.lucene.search.TermsFilter;

public class MySearcher {

public static void main(String[] args) {     String indexPath = "E:\\Lucene\\myindex";   try {    IndexSearcher searcher = new IndexSearcher(indexPath);    String keyword = "的";    Term term = new Term("contents",keyword);    IndexReader indexReader = IndexReader.open(indexPath);    int numberOfDocumentIncludingGivenTerm = indexReader.docFreq(term);    System.out.println("IndexReader的版本为 : "+indexReader.getVersion());    System.out.println("包含词条 ("+term.field()+","+term.text()+") 的Document的数量为 : "+numberOfDocumentIncludingGivenTerm);    Query query = new TermQuery(term);       Date startTime = new Date();       Hits hits = searcher.search(query);    System.out.println("********************************************************************");    int No = 1;    for(int i=0;i<hits.length();i++){     System.out.println("【 序号 】: " + No++);    TermDocs termDocs = searcher.getIndexReader().termDocs(term);     while(termDocs.next()){      if(termDocs.doc() == hits.id(i)){       System.out.println("Document的内部编号为 : "+hits.id(i));       Document doc = hits.doc(i);       List fieldList = doc.getFields();      //System.out.println("==========="+fieldList.size());       System.out.println("Document(编号) "+hits.id(i)+" 的Field的信息: ");       System.out.println("    ------------------------------------");       for(int j=0;j<fieldList.size();j++){        Fieldable field = (Fieldable)fieldList.get(j);        System.out.println("    Field的名称为 : "+field.name());        System.out.println("    Field的内容为 : "+field.stringValue());        System.out.println("    ------------------------------------");       }       System.out.println("Document的内容为 : "+doc);       System.out.println("Document的得分为 : "+hits.score(i));       System.out.println("搜索的该关键字【"+keyword+"】在Document(编号) "+hits.id(i)+" 中,出现过 "+termDocs.freq()+" 次");      }     }     System.out.println("********************************************************************");    }    Date finishTime = new Date();    long timeOfSearch = finishTime.getTime() - startTime.getTime();    System.out.println("本次搜索所用的时间为 "+timeOfSearch+" ms");   } catch (CorruptIndexException e) {    e.printStackTrace();   } catch (IOException e) {    e.printStackTrace();   }}}

首先要保证索引目录E:\\Lucene\\myindex下面已经存在索引文件,可以通过文章 Lucene-2.2.0 源代码阅读学习(4) 中一个使用Lucene的Demo中的递归建立索引的方法,将建立的索引文件存放到E:\\Lucene\\myindex目录之下。

执行上面的主函数,输出结果如下所示:

IndexReader的版本为 : 1207548172961包含词条 (contents,的) 的Document的数量为 : 23********************************************************************【 序号 】: 1Document的内部编号为 : 24Document(编号) 24 的Field的信息:     ------------------------------------    Field的名称为 : path    Field的内容为 : E:\Lucene\txt1\mytxt\FAQ.txt    ------------------------------------    Field的名称为 : modified    Field的内容为 : 200604130754    ------------------------------------Document的内容为 : Document<stored/uncompressed,indexed<path:E:\Lucene\txt1\mytxt\FAQ.txt> stored/uncompressed,indexed<modified:200604130754>>Document的得分为 : 0.5279752搜索的该关键字【的】在Document(编号) 24 中,出现过 291 次********************************************************************【 序号 】: 2Document的内部编号为 : 5Document(编号) 5 的Field的信息:     ------------------------------------    Field的名称为 : path    Field的内容为 : E:\Lucene\txt1\mytxt\3实验题目.txt    ------------------------------------    Field的名称为 : modified    Field的内容为 : 200710300744    ------------------------------------Document的内容为 : Document<stored/uncompressed,indexed<path:E:\Lucene\txt1\mytxt\3实验题目.txt> stored/uncompressed,indexed<modified:200710300744>>Document的得分为 : 0.5252467搜索的该关键字【的】在Document(编号) 5 中,出现过 2 次********************************************************************【 序号 】: 3Document的内部编号为 : 12Document(编号) 12 的Field的信息:     ------------------------------------    Field的名称为 : path    Field的内容为 : E:\Lucene\txt1\mytxt\CustomKeyInfo.txt    ------------------------------------    Field的名称为 : modified    Field的内容为 : 200406041814    ------------------------------------Document的内容为 : Document<stored/uncompressed,indexed<path:E:\Lucene\txt1\mytxt\CustomKeyInfo.txt> stored/uncompressed,indexed<modified:200406041814>>Document的得分为 : 0.51790017搜索的该关键字【的】在Document(编号) 12 中,出现过 70 次********************************************************************【 序号 】: 4Document的内部编号为 : 41Document(编号) 41 的Field的信息:     ------------------------------------    Field的名称为 : path    Field的内容为 : E:\Lucene\txt1\mytxt\Update.txt    ------------------------------------    Field的名称为 : modified    Field的内容为 : 200707050028    ------------------------------------Document的内容为 : Document<stored/uncompressed,indexed<path:E:\Lucene\txt1\mytxt\Update.txt> stored/uncompressed,indexed<modified:200707050028>>Document的得分为 : 0.5059122搜索的该关键字【的】在Document(编号) 41 中,出现过 171 次********************************************************************【 序号 】: 5Document的内部编号为 : 0Document(编号) 0 的Field的信息:     ------------------------------------    Field的名称为 : path    Field的内容为 : E:\Lucene\txt1\120E升级包安装说明.txt    ------------------------------------    Field的名称为 : modified    Field的内容为 : 200803271123    ------------------------------------Document的内容为 : Document<stored/uncompressed,indexed<path:E:\Lucene\txt1\120E升级包安装说明.txt> stored/uncompressed,indexed<modified:200803271123>>Document的得分为 : 0.43770555搜索的该关键字【的】在Document(编号) 0 中,出现过 2 次********************************************************************【 序号 】: 6Document的内部编号为 : 3Document(编号) 3 的Field的信息:     ------------------------------------    Field的名称为 : path    Field的内容为 : E:\Lucene\txt1\mytxt\1实验题目.txt    ------------------------------------    Field的名称为 : modified    Field的内容为 : 200710160733    ------------------------------------Document的内容为 : Document<stored/uncompressed,indexed<path:E:\Lucene\txt1\mytxt\1实验题目.txt> stored/uncompressed,indexed<modified:200710160733>>Document的得分为 : 0.4333064搜索的该关键字【的】在Document(编号) 3 中,出现过 1 次********************************************************************【 序号 】: 7Document的内部编号为 : 60Document(编号) 60 的Field的信息:     ------------------------------------    Field的名称为 : path    Field的内容为 : E:\Lucene\txt1\mytxt\猫吉又有个忙,需要大家帮忙一下.txt    ------------------------------------    Field的名称为 : modified    Field的内容为 : 200706161112    ------------------------------------Document的内容为 : Document<stored/uncompressed,indexed<path:E:\Lucene\txt1\mytxt\猫吉又有个忙,需要大家帮忙一下.txt> stored/uncompressed,indexed<modified:200706161112>>Document的得分为 : 0.4106042搜索的该关键字【的】在Document(编号) 60 中,出现过 11 次********************************************************************【 序号 】: 8Document的内部编号为 : 59Document(编号) 59 的Field的信息:     ------------------------------------    Field的名称为 : path    Field的内容为 : E:\Lucene\txt1\mytxt\汉化说明.txt    ------------------------------------    Field的名称为 : modified    Field的内容为 : 200708210247    ------------------------------------Document的内容为 : Document<stored/uncompressed,indexed<path:E:\Lucene\txt1\mytxt\汉化说明.txt> stored/uncompressed,indexed<modified:200708210247>>Document的得分为 : 0.39057708搜索的该关键字【的】在Document(编号) 59 中,出现过 13 次********************************************************************【 序号 】: 9Document的内部编号为 : 44Document(编号) 44 的Field的信息:     ------------------------------------    Field的名称为 : path    Field的内容为 : E:\Lucene\txt1\mytxt\Visual Studio 2005注册升级.txt    ------------------------------------    Field的名称为 : modified    Field的内容为 : 200801300512    ------------------------------------Document的内容为 : Document<stored/uncompressed,indexed<path:E:\Lucene\txt1\mytxt\Visual Studio 2005注册升级.txt> stored/uncompressed,indexed<modified:200801300512>>Document的得分为 : 0.37525433搜索的该关键字【的】在Document(编号) 44 中,出现过 3 次********************************************************************【 序号 】: 10Document的内部编号为 : 56Document(编号) 56 的Field的信息:     ------------------------------------    Field的名称为 : path    Field的内容为 : E:\Lucene\txt1\mytxt\新1建 文本文档.txt    ------------------------------------    Field的名称为 : modified    Field的内容为 : 200710311142    ------------------------------------Document的内容为 : Document<stored/uncompressed,indexed<path:E:\Lucene\txt1\mytxt\新1建 文本文档.txt> stored/uncompressed,indexed<modified:200710311142>>Document的得分为 : 0.36621076搜索的该关键字【的】在Document(编号) 56 中,出现过 35 次********************************************************************【 序号 】: 11Document的内部编号为 : 46Document(编号) 46 的Field的信息:     ------------------------------------    Field的名称为 : path    Field的内容为 : E:\Lucene\txt1\mytxt\使用技巧集萃.txt    ------------------------------------    Field的名称为 : modified    Field的内容为 : 200511210413    ------------------------------------Document的内容为 : Document<stored/uncompressed,indexed<path:E:\Lucene\txt1\mytxt\使用技巧集萃.txt> stored/uncompressed,indexed<modified:200511210413>>Document的得分为 : 0.35693806搜索的该关键字【的】在Document(编号) 46 中,出现过 133 次********************************************************************【 序号 】: 12Document的内部编号为 : 30Document(编号) 30 的Field的信息:     ------------------------------------    Field的名称为 : path    Field的内容为 : E:\Lucene\txt1\mytxt\MyEclipse 注册码.txt    ------------------------------------    Field的名称为 : modified    Field的内容为 : 200712061059    ------------------------------------Document的内容为 : Document<stored/uncompressed,indexed<path:E:\Lucene\txt1\mytxt\MyEclipse 注册码.txt> stored/uncompressed,indexed<modified:200712061059>>Document的得分为 : 0.3460366搜索的该关键字【的】在Document(编号) 30 中,出现过 5 次********************************************************************【 序号 】: 13Document的内部编号为 : 63Document(编号) 63 的Field的信息:     ------------------------------------    Field的名称为 : path    Field的内容为 : E:\Lucene\txt1\疑问即时记录.txt    ------------------------------------    Field的名称为 : modified    Field的内容为 : 200711141408    ------------------------------------Document的内容为 : Document<stored/uncompressed,indexed<path:E:\Lucene\txt1\疑问即时记录.txt> stored/uncompressed,indexed<modified:200711141408>>Document的得分为 : 0.30325133搜索的该关键字【的】在Document(编号) 63 中,出现过 6 次********************************************************************【 序号 】: 14Document的内部编号为 : 37Document(编号) 37 的Field的信息:     ------------------------------------    Field的名称为 : path    Field的内容为 : E:\Lucene\txt1\mytxt\readme.txt    ------------------------------------    Field的名称为 : modified    Field的内容为 : 200803101314    ------------------------------------Document的内容为 : Document<stored/uncompressed,indexed<path:E:\Lucene\txt1\mytxt\readme.txt> stored/uncompressed,indexed<modified:200803101314>>Document的得分为 : 0.26262334搜索的该关键字【的】在Document(编号) 37 中,出现过 8 次********************************************************************【 序号 】: 15Document的内部编号为 : 48Document(编号) 48 的Field的信息:     ------------------------------------    Field的名称为 : path    Field的内容为 : E:\Lucene\txt1\mytxt\剑心补丁使用说明(readme).txt    ------------------------------------    Field的名称为 : modified    Field的内容为 : 200803101357    ------------------------------------Document的内容为 : Document<stored/uncompressed,indexed<path:E:\Lucene\txt1\mytxt\剑心补丁使用说明(readme).txt> stored/uncompressed,indexed<modified:200803101357>>Document的得分为 : 0.26262334搜索的该关键字【的】在Document(编号) 48 中,出现过 8 次********************************************************************【 序号 】: 16Document的内部编号为 : 47Document(编号) 47 的Field的信息:     ------------------------------------    Field的名称为 : path    Field的内容为 : E:\Lucene\txt1\mytxt\关系记录.txt    ------------------------------------    Field的名称为 : modified    Field的内容为 : 200802201145    ------------------------------------Document的内容为 : Document<stored/uncompressed,indexed<path:E:\Lucene\txt1\mytxt\关系记录.txt> stored/uncompressed,indexed<modified:200802201145>>Document的得分为 : 0.23161201搜索的该关键字【的】在Document(编号) 47 中,出现过 14 次********************************************************************【 序号 】: 17Document的内部编号为 : 40Document(编号) 40 的Field的信息:     ------------------------------------    Field的名称为 : path    Field的内容为 : E:\Lucene\txt1\mytxt\Struts之AddressBooks学习笔记.txt    ------------------------------------    Field的名称为 : modified    Field的内容为 : 200710131711    ------------------------------------Document的内容为 : Document<stored/uncompressed,indexed<path:E:\Lucene\txt1\mytxt\Struts之AddressBooks学习笔记.txt> stored/uncompressed,indexed<modified:200710131711>>Document的得分为 : 0.21885277搜索的该关键字【的】在Document(编号) 40 中,出现过 8 次********************************************************************【 序号 】: 18Document的内部编号为 : 51Document(编号) 51 的Field的信息:     ------------------------------------    Field的名称为 : path    Field的内容为 : E:\Lucene\txt1\mytxt\密码强度检验.txt    ------------------------------------    Field的名称为 : modified    Field的内容为 : 200712010901    ------------------------------------Document的内容为 : Document<stored/uncompressed,indexed<path:E:\Lucene\txt1\mytxt\密码强度检验.txt> stored/uncompressed,indexed<modified:200712010901>>Document的得分为 : 0.12380183搜索的该关键字【的】在Document(编号) 51 中,出现过 1 次********************************************************************【 序号 】: 19Document的内部编号为 : 50Document(编号) 50 的Field的信息:     ------------------------------------    Field的名称为 : path    Field的内容为 : E:\Lucene\txt1\mytxt\史上最强天籁之声及欧美流行曲超级精选【 FLAC分轨】.txt    ------------------------------------    Field的名称为 : modified    Field的内容为 : 200712231241    ------------------------------------Document的内容为 : Document<stored/uncompressed,indexed<path:E:\Lucene\txt1\mytxt\史上最强天籁之声及欧美流行曲超级精选【 FLAC分轨】.txt> stored/uncompressed,indexed<modified:200712231241>>Document的得分为 : 0.1083266搜索的该关键字【的】在Document(编号) 50 中,出现过 1 次********************************************************************【 序号 】: 20Document的内部编号为 : 57Document(编号) 57 的Field的信息:     ------------------------------------    Field的名称为 : path    Field的内容为 : E:\Lucene\txt1\mytxt\新建 文本文档.txt    ------------------------------------    Field的名称为 : modified    Field的内容为 : 200710270258    ------------------------------------Document的内容为 : Document<stored/uncompressed,indexed<path:E:\Lucene\txt1\mytxt\新建 文本文档.txt> stored/uncompressed,indexed<modified:200710270258>>Document的得分为 : 0.09285137搜索的该关键字【的】在Document(编号) 57 中,出现过 4 次********************************************************************【 序号 】: 21Document的内部编号为 : 45Document(编号) 45 的Field的信息:     ------------------------------------    Field的名称为 : path    Field的内容为 : E:\Lucene\txt1\mytxt\书籍网站.txt    ------------------------------------    Field的名称为 : modified    Field的内容为 : 200708071255    ------------------------------------Document的内容为 : Document<stored/uncompressed,indexed<path:E:\Lucene\txt1\mytxt\书籍网站.txt> stored/uncompressed,indexed<modified:200708071255>>Document的得分为 : 0.0670097搜索的该关键字【的】在Document(编号) 45 中,出现过 3 次********************************************************************【 序号 】: 22Document的内部编号为 : 61Document(编号) 61 的Field的信息:     ------------------------------------    Field的名称为 : path    Field的内容为 : E:\Lucene\txt1\mytxt\网络查询大全.txt    ------------------------------------    Field的名称为 : modified    Field的内容为 : 200111200655    ------------------------------------Document的内容为 : Document<stored/uncompressed,indexed<path:E:\Lucene\txt1\mytxt\网络查询大全.txt> stored/uncompressed,indexed<modified:200111200655>>Document的得分为 : 0.065655835搜索的该关键字【的】在Document(编号) 61 中,出现过 2 次********************************************************************【 序号 】: 23Document的内部编号为 : 14Document(编号) 14 的Field的信息:     ------------------------------------    Field的名称为 : path    Field的内容为 : E:\Lucene\txt1\mytxt\CustomKeysSample.txt    ------------------------------------    Field的名称为 : modified    Field的内容为 : 200610100451    ------------------------------------Document的内容为 : Document<stored/uncompressed,indexed<path:E:\Lucene\txt1\mytxt\CustomKeysSample.txt> stored/uncompressed,indexed<modified:200610100451>>Document的得分为 : 0.051179506搜索的该关键字【的】在Document(编号) 14 中,出现过 7 次********************************************************************本次搜索所用的时间为 187 ms

 

其中,IndexReader是一个用于读取索引文件的抽象类,可以参考该类的源代码。IndexReader类可以很方便地打开一个索引目录(即创建一个输入流),这要使用到它的静态(static)方法open()打开即可,然后就可以访问索引文件了,从而实现对索引文件的维护。

IndexReader类实现了各种打开索引文件的方式,由于定义为static的,所以非常方便地调用,如下所示:

public static IndexReader open(String path) throws CorruptIndexException, IOException {    // 通过String的索引目录的路径    return open(FSDirectory.getDirectory(path), true, null);}

public static IndexReader open(File path) throws CorruptIndexException, IOException {    // 通过File构造的索引目录文件    return open(FSDirectory.getDirectory(path), true, null);}

public static IndexReader open(final Directory directory) throws CorruptIndexException, IOException { // 直接通过Directory来打开    return open(directory, false, null);}

public static IndexReader open(final Directory directory, IndexDeletionPolicy deletionPolicy) throws CorruptIndexException, IOException { // 直接通过Directory来打开,并指定一种索引文件删除策略,可以对索引文件进行维护(删除操作)    return open(directory, false, deletionPolicy);}

其中,最核心的实现是在一个私有的open()方法中实现的,如下所示:

private static IndexReader open(final Directory directory, final boolean closeDirectory, final IndexDeletionPolicy deletionPolicy) throws CorruptIndexException, IOException {

    return (IndexReader) new SegmentInfos.FindSegmentsFile(directory) {

      protected Object doBody(String segmentFileName) throws CorruptIndexException, IOException {

        SegmentInfos infos = new SegmentInfos();        infos.read(directory, segmentFileName);

        IndexReader reader;

        if (infos.size() == 1) {    // index is optimized          reader = SegmentReader.get(infos, infos.info(0), closeDirectory);        } else {

         // To reduce the chance of hitting FileNotFound          // (and having to retry), we open segments in          // reverse because IndexWriter merges & deletes          // the newest segments first.

          IndexReader[] readers = new IndexReader[infos.size()];          for (int i = infos.size()-1; i >= 0; i--) {            try {              readers[i] = SegmentReader.get(infos.info(i));            } catch (IOException e) {              // Close all readers we had opened:              for(i++;i<infos.size();i++) {                readers[i].close();              }              throw e;            }          }

          reader = new MultiReader(directory, infos, closeDirectory, readers);        }        reader.deletionPolicy = deletionPolicy;        return reader;      }    }.run();}

上面测试程序中,IndexSearcher类是实现检索的核心类。它提供了很多中不同的检索方式,返回的对象也可以适用于不同的需要,比如Hits、TopFieldDocs、TopDocs,而且,还可以指定排序Sort、权重Weight、过滤器Filter作为search()方法的参数,用起来的灵活、方便。

通过程序中,红色标注的代码行:

TermDocs termDocs = searcher.getIndexReader().termDocs(term);

其实,一个IndexSearcher实例化以后,可以通过它获取到一个IndexReader的实例,从而打开一个索引目录。

然后从就可以从创建的输入流中读取索引文件的详细信息:

1、每个Document的内部编号(是唯一的,可以通过这个编号对其进行维护);

2、每个Document中都有多个Field,可以读取Field的名称、路径、Field的内容等等。

上面的测试程序中,没有输出名称为“contents”的Field,是因为在索引文件中没有存储Fielde的内容(即文本信息)。因为Field的内容是根据从指定的数据源中获取,而数据源可能是数据量非常大的一些文件,如果直接将它们保存到索引文件中,会占用很大的磁盘空间。

其实,可以根据需要存储。上面之所以没有存储,可以追溯到Lucene自带的Demo中的设置,在org.apache.lucene.demo.FileDocument中创建Field,如下所示:

    // 构造一个Field,这个Field可以从一个文件流中读取,必须保证由f所构造的文件流是打开的    doc.add(new Field("contents", new FileReader(f)));

然后,看Field的该构造方法的定义:

public Field(String name, Reader reader) {    this(name, reader, TermVector.NO);}

这个构造方法指定了要为这个创建的Field进行分词、索引,但是不存储。

可以参考调用的另一个构造方法:

public Field(String name, Reader reader, TermVector termVector) {    if (name == null)      throw new NullPointerException("name cannot be null");    if (reader == null)      throw new NullPointerException("reader cannot be null");        this.name = name.intern();        // field names are interned    this.fieldsData = reader;        this.isStored = false;    // 指定不进行存储    this.isCompressed = false;        this.isIndexed = true;    // 要进行索引    this.isTokenized = true;    // 要进行分词        this.isBinary = false;        setStoreTermVector(termVector);}

上面测试程序中,Hits是一个内容非常丰富的实现类。

通过检索返回Hits类的对象,可以通过Hits类的实例来获取Document的id,以及计算该Document的得分,并且在search()执行之后,返回的结果按照得分的高低来排序输出,得分值高的在前面。

以此做个引子,之后再详细学习研究。

 

 

 

转载请注明原文地址: https://www.6miu.com/read-5052690.html

最新回复(0)