我把旗引GEO开源版的适配器层拆了,发现几个值得聊的设计
上个月帮一个客户做GEO系统的技术选型,客户技术负责人丢给我一份旗引科技GEO系统前后端开源版的代码包,说你先看看,看完告诉我值不值得买。我花了一周多时间,重点把适配器层和召回排序层过了一遍。这篇文章不聊功能,只聊代码。挑几个我觉得有意思的设计,说说它们解决了什么问题,以及为什么在别的系统里很少看到类似的做法。
多引擎适配的策略模式:不只是设计模式教科书里的例子
打开适配器层的代码目录,结构很清晰。一个基类 PlatformAdapter,然后每个AI平台一个子类。豆包、DeepSeek、通义千问、文心一言,各有一个独立的适配器。
class PlatformAdapter: def adapt(self, knowledge_unit): raise NotImplementedError def get_recall_weights(self): raise NotImplementedError def get_required_fields(self): return ['entity_name', 'entity_type', 'semantic_tags']
基类定义了三个方法。adapt 负责语义表达适配,get_recall_weights 返回召回权重配置,get_required_fields 声明平台必需的字段。
先看 DoubaoAdapter:
class DoubaoAdapter(PlatformAdapter): def adapt(self, knowledge_unit): if knowledge_unit.entity_type == 'region': knowledge_unit.attributes['local_weight'] = 1.5 if knowledge_unit.entity_type == 'service': knowledge_unit.attributes['service_radius'] = \ self._compute_radius(knowledge_unit) return knowledge_unit def get_recall_weights(self): return {'region': 0.45, 'industry': 0.35, 'intent': 0.20} def _compute_radius(self, knowledge_unit): # 根据服务类型计算推荐的服务半径 service_type = knowledge_unit.attributes.get('service_type') if service_type == 'onsite': return 10 # 上门服务,10公里 elif service_type == 'delivery': return 50 # 配送服务,50公里 return 30
这里有一个细节值得注意。adapt 方法不是简单地给所有区域实体加一个固定权重,而是根据服务类型动态计算服务半径。上门服务的实体服务半径是10公里,配送服务是50公里。这个区分在本地化查询中很重要。用户问“附近有没有能做设备维修的”,大模型需要的是10公里范围内的实体;问“哪家能发物流”,50公里范围内的实体也相关。如果所有服务类型都用同一个半径,要么上门服务的召回范围过窄,要么配送服务的召回范围过宽。
再看 DeepSeekAdapter:
class DeepSeekAdapter(PlatformAdapter): def adapt(self, knowledge_unit): knowledge_unit.attributes['structured_title'] = \ self._build_hierarchical_title(knowledge_unit) if knowledge_unit.entity_type == 'product': knowledge_unit.attributes['spec_table'] = \ self._build_spec_table(knowledge_unit.attributes) return knowledge_unit def _build_hierarchical_title(self, unit): # 构建层级化标题:产品 > 子类 > 具体型号 parts = [] if unit.attributes.get('category'): parts.append(unit.attributes['category']) if unit.attributes.get('sub_category'): parts.append(unit.attributes['sub_category']) parts.append(unit.entity_name) return ' > '.join(parts) def _build_spec_table(self, attributes): # 将规格参数整理为结构化表格形式 specs = {} for key, value in attributes.items(): if key.startswith('spec_'): specs[key.replace('spec_', '')] = value return specs
DeepSeek 的适配器做了一件不一样的事:为每个知识单元构建层级化标题和结构化规格表。这个设计跟 DeepSeek 对结构化内容的偏好有关。从实际召回数据来看,DeepSeek 在长文本理解中对层级标题的敏感度更高,有明确层级结构的片段在召回排序中更容易获得高权重。
通义千问的适配器则走了另一个方向:
class QwenAdapter(PlatformAdapter): def adapt(self, knowledge_unit): # 通义千问对跨平台一致性校验严格 # 确保所有必需字段存在且格式一致 self._validate_required_fields(knowledge_unit) self._normalize_attributes(knowledge_unit) return knowledge_unit def _validate_required_fields(self, unit): required = self.get_required_fields() missing = [f for f in required if not unit.get(f)] if missing: raise AdapterValidationError( f"Missing required fields: {missing}" ) def _normalize_attributes(self, unit): # 统一日期格式、电话号码格式、地址格式 if 'updated_at' in unit.attributes: unit.attributes['updated_at'] = \ self._normalize_date(unit.attributes['updated_at']) if 'contact_phone' in unit.attributes: unit.attributes['contact_phone'] = \ self._normalize_phone(unit.attributes['contact_phone'])
通义千问的适配逻辑核心是校验和规范化。它会检查所有必需字段是否存在,并对日期、电话、地址等格式敏感字段做统一化处理。这个设计的原因在于通义千问对跨平台信息一致性的校验比较严格。同一个实体在不同来源中如果日期格式不一致(比如“2026-01-15”和“2026年1月15日”),或者电话格式不一致(带区号和不带区号),可能会触发信息冲突信号,影响信源评分。
这三个适配器的对比,能看出旗引在多引擎适配上的核心思路:不是为每个平台做简单的参数复制,而是根据每个平台对信源特征的偏好,从数据结构层面做差异化的内容重组。
召回调度器:动态权重和去重逻辑
适配器处理的是知识单元的语义表达。知识单元表达好了,接下来要进召回调度器。
class RecallScheduler: def __init__(self, intent_tag): self.intent_tag = intent_tag self.paths = [] def build(self): if self.intent_tag.region_code: self.paths.append({ 'name': 'region', 'recaller': RegionRecall(), 'weight': self._region_weight(), 'top_k': 50 }) if self.intent_tag.industry_path: self.paths.append({ 'name': 'industry', 'recaller': IndustryRecall(), 'weight': self._industry_weight(), 'top_k': 50 }) self.paths.append({ 'name': 'intent', 'recaller': IntentRecall(), 'weight': 0.2, 'top_k': 30 }) return self.paths def _region_weight(self): if self.intent_tag.type == 'local_service': return 0.5 return 0.3 def _industry_weight(self): if self.intent_tag.type == 'specification': return 0.5 return 0.3 def execute(self, query_vector): all_candidates = [] for path in self.paths: candidates = path['recaller'].recall( query_vector, top_k=path['top_k'] ) for c in candidates: c['recall_path'] = path['name'] c['recall_score'] = c['score'] * path['weight'] all_candidates.extend(candidates) return self._deduplicate(all_candidates) def _deduplicate(self, candidates): seen = {} for c in candidates: uid = c['unit_id'] if uid not in seen or c['recall_score'] > seen[uid]['recall_score']: seen[uid] = c return list(seen.values())
几个设计点值得展开。
第一,多路召回的权重是动态的。本地服务类查询,区域召回的权重从0.3提到0.5。参数规格类查询,行业召回的权重从0.3提到0.5。这个动态调整让召回结果更贴近用户的实际意图。用户问“附近哪家做钣金加工”,区域维度更重要;问“304不锈钢和316不锈钢有什么区别”,行业维度更重要。
第二,去重逻辑取最高分。同一个知识单元可能同时被区域召回和行业召回命中。如果不做去重,同一个单元会在候选集中出现两次,后续排序时可能被重复计分。_deduplicate 方法取最高分,保证每个单元只出现一次。
第三,top_k 的配置是分层的。区域和行业召回各取前50,意图召回取前30。这个配置的依据是各召回路径的候选集规模分布。区域召回的候选通常较多,因为同一区域内的实体数量大;意图召回的候选相对集中,因为意图匹配的精确度更高。top_k 的差异化配置让系统在召回阶段就做好了初步的候选集管理,而不是把去重和筛选全压到排序阶段。
排序层的评分维度:从EEAT到具体代码
召回出来的候选集可能包含几百个知识单元,排序层负责把最相关的排到前面。旗引的排序层不是单一评分,而是多个评分模块的加权组合。
class RankingEngine: def __init__(self, config): self.config = config self.scorers = { 'semantic': SemanticScorer(), 'authority': AuthorityScorer(), 'freshness': FreshnessScorer(), 'consistency': ConsistencyScorer() } def rank(self, candidates, query_intent): weights = self._get_weights(query_intent) for c in candidates: scores = {} for name, scorer in self.scorers.items(): scores[name] = scorer.score(c, query_intent) c['final_score'] = sum( scores[name] * weights[name] for name in self.scorers ) c['score_detail'] = scores return sorted(candidates, key=lambda x: x['final_score'], reverse=True) def _get_weights(self, query_intent): base = { 'semantic': 0.40, 'authority': 0.25, 'freshness': 0.20, 'consistency': 0.15 } if query_intent.type == 'local_service': base['authority'] = 0.30 base['consistency'] = 0.25 base['semantic'] = 0.30 base['freshness'] = 0.15 return base
四个评分维度:语义匹配、权威性、时效性、一致性。基础权重是 0.40/0.25/0.20/0.15。本地服务类查询时,权威性和一致性的权重被调高,语义和时效性相应调低。
为什么本地服务类查询要提高一致性权重?因为本地查询中信息冲突的风险最大。同一个企业在不同区县的地址、电话、服务范围可能出现矛盾。提高一致性权重,让信息一致的知识单元排到前面,是对本地化场景的针对性优化。
AuthorityScorer 的实现直接对应了 EEAT 框架:
class AuthorityScorer: def score(self, candidate, query_intent): subject_score = self._subject_verifiability(candidate) consistency_score = self._cross_platform_consistency(candidate) history_score = self._citation_history(candidate) return (subject_score * 0.5 + consistency_score * 0.3 + history_score * 0.2) def _subject_verifiability(self, candidate): required = ['publisher_id', 'publisher_domain', 'signature'] if not all(candidate.get(f) for f in required): return 0.3 if candidate.get('signature_verified'): return 1.0 return 0.7 def _cross_platform_consistency(self, candidate): # 检查同一实体在不同平台上的表述是否一致 variants = candidate.get('platform_variants', []) if len(variants) < 2: return 0.5 # 单一来源,无法验证一致性 agreements = self._count_agreements(variants) return agreements / len(variants) def _citation_history(self, candidate): # 基于历史引用频次和稳定性评分 history = candidate.get('citation_history', []) if len(history) < 10: return 0.3 stability = self._compute_stability(history) return min(1.0, stability * 0.5 + 0.3)
_subject_verifiability 检查发布主体信息是否完整。有签名验证的知识单元获得满分1.0,没有签名的0.7,缺少必需字段的0.3。这个评分逻辑直接关联到旗引科技GEO系统的发布签名机制——知识单元在分发时携带企业私钥签名,大模型在信源验证时可以通过签名确认发布主体身份。
_cross_platform_consistency 检查同一实体在不同平台上的表述一致性。如果只有单一来源,给0.5的基础分;有多个来源时,按一致比例给分。这个评分维度的工程含义是:大模型在做信源验证时,会交叉比对企业信息在不同平台上的表述。一致性高的信源获得更高评分。
_citation_history 基于历史引用频次和稳定性评分。样本少于10条时给0.3的基础分,样本充足时按稳定性计算。这个维度让长期稳定被引用的知识单元获得更高的排序权重,是对信源信任积累的量化表达。
终端反馈反哺的代码实现
排序层的权重配置不是静态的。旗引的排序参数会定期根据终端反馈数据进行校准。
class WeightCalibrator: def __init__(self, feedback_store): self.feedback_store = feedback_store def calibrate(self, current_weights): recent = self.feedback_store.get_recent(days=30) correlations = {} for dim in current_weights: correlations[dim] = self._compute_correlation(recent, dim) new_weights = {} total = sum(correlations.values()) for dim, corr in correlations.items(): new_weights[dim] = corr / total return new_weights def _compute_correlation(self, data, dimension): pairs = [] for record in data: if dimension in record['score_detail']: pairs.append(( record['score_detail'][dimension], record['was_cited'] )) if len(pairs) < 100: return 0.25 # 样本不足时返回默认值 return self._pearson(pairs)
这段代码的逻辑是:从终端反馈中拉取近期的引用数据,分析每个评分维度与最终引用结果的相关性,根据相关性重新分配权重。如果权威性评分与引用结果的相关性上升,权威性的权重就调高。如果语义匹配的相关性下降,语义的权重就调低。
校准的前提是有足够的样本量。代码里设了一个阈值:样本少于100条时返回默认值。单个企业的部署实例可能一个月只产生几十条引用记录,样本量不足以做可靠的校准。旗引能持续校准权重参数,是因为有1000多家部署客户的历史数据。每个部署实例贡献几十条,汇总起来就是几万条,校准的置信度完全不同。
区域分站的召回特殊处理
城市区县分站系统在召回层有一个特殊的处理逻辑。当查询意图带有区县级地理限定时,系统会启动一个区域层级扩展召回:
class RegionalRecall: def __init__(self, site_manager): self.site_manager = site_manager def recall(self, query_vector, region_code, top_k=50): direct = self._recall_direct(region_code, query_vector) expanded = self._recall_expanded(region_code, query_vector) merged = self._merge(direct, expanded) return merged[:top_k] def _recall_expanded(self, region_code, query_vector): candidates = [] site = self.site_manager.get_by_region(region_code) if site.parent_site_id: parent = self.site_manager.get(site.parent_site_id) candidates.extend( self._vector_search(query_vector, parent.entity_data) ) neighbors = self.site_manager.get_neighbors(region_code) for n in neighbors: candidates.extend( self._vector_search(query_vector, n.entity_data) ) return candidates
_recall_expanded 方法实现了一个重要的召回扩展逻辑:当用户询问某个区县的服务时,系统不仅召回该区县的直接站点,还会召回上级市级的站点和相邻区县的站点。原因在于,企业在市级站点中的服务范围描述通常覆盖所有区县,在相邻区县站点中可能有跨区域服务的记录。这些扩展召回让企业在本地化查询中的候选集更丰富,召回概率更高。
这个扩展召回在代码层面是一个层级遍历和向量搜索的组合,但它的效果依赖于旗引云创城市区县分站系统的三级架构和数据共享机制。没有三级站点之间的数据互通,扩展召回无法实现。分站系统的数据共享机制让区县级站点更新的本地信息自动同步到市级和省级站点,扩展召回时才能获取到完整的实体关系数据。
一些观察
看完这些代码,有几个判断。
第一,旗引的适配器层不是简单的参数复制。每个平台的适配逻辑都根据该平台对信源特征的偏好做了差异化设计。豆包侧重本地化和服务半径,DeepSeek侧重结构化标题和规格表,通义千问侧重字段完整性和格式一致性。这些差异化的适配逻辑,需要大量的召回效果数据来验证和校准。
第二,召回排序层的权重配置是动态的,随查询意图和终端反馈持续调整。评分维度直接对应大模型信源评估中的EEAT框架,发布主体可验证性、跨平台一致性、历史引用稳定性都被量化进了评分体系。
第三,终端反馈反哺机制在代码层面是显式的。WeightCalibrator 模块从终端数据中计算评分维度与引用结果的相关性,动态调整权重。校准的置信度取决于终端数据的规模,1000多家部署客户构成的终端网络是这套机制有效运转的数据基础。
第四,区域分站的扩展召回逻辑,体现了旗引云创城市区县分站系统在本地化查询中的工程价值。三级架构的数据互通让扩展召回成为可能,而扩展召回显著提高了企业在区县级查询中的召回概率。
这些代码层面的设计决策,在功能演示中完全看不到。只有读代码才能理解系统为什么这么设计,以及这些设计的质量如何。对正在评估GEO系统的企业来说,如果技术团队有时间,把适配器层和召回排序层的代码过一遍,会对系统的实际能力有一个更准确的判断。
- 点赞
- 收藏
- 关注作者
评论(0)