ARTICLE DETAIL

资讯详情

深耕网站视觉设计与运营推广的一线实战洞察。

Magika standard_v3_0 模型输出内容类型全量清单解析:213 个 Content Type Label 与识别原理

Magika standard_v3_0 模型输出内容类型全量清单解析:213 个 Content Type Label 与识别原理 Magika standard_v3_0 模型输出内容类型全量清单解析213 个 Content Type Label 与识别原理【免费下载链接】magikaFast and accurate AI powered file content types detection项目地址: https://gitcode.com/GitHub_Trending/ma/magika本文围绕 Magika 当前默认深度学习模型standard_v3_0的输出标签空间展开完整列出模型可能输出的 213 个内容类型标签Content Type Label并结合该模型自带的 config.min.json 与 Python 端推理实现 magika.py讲清标签的生成链路、预测模式、置信度阈值与输出结构。读完你将能准确解读 Magika 的任意识别结果并能在自己的代码中直接使用get_supported_content_types()查询当前模型的全部可输出类型。一、standard_v3_0当前默认模型的输出标签全集Magika 在 Python 端将standard_v3_0设为默认模型见 magika.py 中的DEFAULT_MODEL_NAME standard_v3_0模型文件存放在 assets/models/standard_v3_0/。该模型的配置 config.min.json 中target_labels_space字段定义了模型输出层的完整标签空间共 213 个标签与模型自带 READMEassets/models/standard_v3_0/README.md列出的清单一一对应。与上一代standard_v2_1相比v3.0 的标签空间在末尾新增了randombytes、randomtxt、symlinktext三个内部标签见 standard_v2_1/config.min.json 中仅有unknown的对比并通过overwrite_map将它们映射回用户可见的常规标签详见下文“覆盖映射”小节同时为handlebars、markdown增加了专属高置信度阈值。1.1 完整标签清单按索引 1–106下表完整继承自模型 README索引、标签名Content Type Label与说明Description均保持原样IndexContent Type LabelDescription13gp3GPP multimedia file2aceACE archive3aiAdobe Illustrator Artwork4aidlAndroid Interface Definition Language5apkAndroid package6applebplistApple binary property list7appleplistApple property list8asmAssembly9aspASP source10autohotkeyAutoHotKey script11autoitAutoIt script12awkAwk13batchDOS batch file14bazelBazel build file15bibBibTeX16bmpBMP image data17bzipbzip2 compressed data18cC source19cabMicrosoft Cabinet archive data20catWindows Catalog file21chmMS Windows HtmlHelp Data22clojureClojure23cmakeCMake build file24cobolCobol25coffIntel 80386 COFF26coffeescriptCoffeeScript27cppC source28crtCertificates (binary format)29crxGoogle Chrome extension30csC# source31csproj.NET project config32cssCSS source33csvCSV document34dartDart source35debDebian binary package36dexDalvik dex file37dicomDICOM38diffDiff file39dmDream Maker40dmgApple disk image41docMicrosoft Word CDF document42dockerfileDockerfile43docxMicrosoft Word 2007 document44dsstoreApplication Desktop Services Store45dwgAutocad Drawing46dxfAudocad Drawing Exchange Format47elfELF executable48elixirElixir script49emfWindows Enhanced Metafile image data50emlRFC 822 mail51epubEPUB document52erbEmbedded Ruby source53erlangErlang source54flacFLAC audio bitstream data55flvFlash Video56fortranFortran57gemfileGemfile file58gemspecGemspec file59gifGIF image data60gitattributesGitattributes file61gitmodulesGitmodules file62goGolang source63gradleGradle source64groovyGroovy source65gzipgzip compressed data66h5Hierarchical Data Format v567handlebarsHandlebars source68haskellHaskell source69hclHashiCorp configuration language70hlpMS Windows help71htaccessApache access configuration72htmlHTML document73icnsMac OS X icon74icoMS Windows icon resource75icsInternet Calendaring and Scheduling76ignorefileIgnorefile77iniINI configuration file78internetshortcutMS Windows Internet shortcut79ipynbJupyter notebook80isoISO 9660 CD-ROM filesystem data81jarJava archive data (JAR)82javaJava source83javabytecodeJava compiled bytecode84javascriptJavaScript source85jinjaJinja template86jp2jpeg200087jpegJPEG image data88jsonJSON document89jsonlJSONL document90juliaJulia source91kotlinKotlin source92latexLaTeX document93lhaLHarc archive94lispLisp source95lnkMS Windows shortcut96luaLua97m3uM3U playlist98m4GNU Macro99machoMach-O executable100makefileMakefile source101markdownMarkdown document102matlabMatlab Source103mhtMHTML document104midiMidi105mkvMatroska106mp3MP3 media file1.2 完整标签清单按索引 107–213IndexContent Type LabelDescription107mp4MP4 media file108mscompressMS Compress archive data109msiMicrosoft Installer file110mumWindows Update Package file111npyNumpy Array112npzNumpy Arrays Archive113nupkgNuGet Package114objectivecObjectiveC source115ocamlOCaml116odpOpenDocument Presentation117odsOpenDocument Spreadsheet118odtOpenDocument Text119oggOgg data120oneOne Note121onnxOpen Neural Network Exchange122otfOpenType font123outlookMS Outlook Message124parquetApache Parquet125pascalPascal source126pcappcap capture file127pdbWindows Program Database128pdfPDF document129pebinPE Windows executable130pemPEM certificate131perlPerl source132phpPHP source133picklePython pickle134pngPNG image135poPortable Object (PO) for i18n136postscriptPostScript document137powershellPowershell source138pptMicrosoft PowerPoint CDF document139pptxMicrosoft PowerPoint 2007 document140prologProlog source141proteindbProtein DB142protoProtocol buffer definition143psdAdobe Photoshop144pythonPython source145pythonbytecodePython compiled bytecode146pytorchPytorch storage file147qtQuickTime148rR (language)149rarRAR archive data150rdfResource Description Framework document (RDF)151rpmRedHat Package Manager archive (RPM)152rstReStructuredText document153rtfRich Text Format document154rubyRuby source155rustRust source156scalaScala source157scssSCSS source158sevenzip7-zip archive data159sgmlsgml160shellShell script161smaliSmali source162snapSnap archive163soliditySolidity source164sqlSQL source165sqliteSQLITE database166squashfsSquash filesystem167srtSubRip Text Format168stlbinaryStereolithography CAD (binary)169stltextStereolithography CAD (text)170sumChecksum file171svgSVG Scalable Vector Graphics image data172swfSmall Web File173swiftSwift174tarPOSIX tar archive175tclTickle176textprotoText protocol buffer177tgaTarga image data178thumbsdbWindows thumbnail cache179tiffTIFF image data180tomlToms obvious, minimal language181torrentBitTorrent file182tsvTSV document183ttfTrueType Font data184twigTwig template185txtGeneric text document186typescriptTypescript187unknownUnknown binary data188vbaMS Visual Basic source (VBA)189vcxprojVisual Studio MSBuild project190verilogVerilog source191vhdlVHDL source192vttWeb Video Text Tracks193vueVue source194wasmWeb Assembly195wavWaveform Audio file (WAV)196webmWebM media file197webpWebP media file198winregistryWindows Registry text199wmfWindows metafile200woffWeb Open Font Format201woff2Web Open Font Format v2202xarXAR archive compressed data203xlsMicrosoft Excel CDF document204xlsbMicrosoft Excel 2007 document (binary format)205xlsxMicrosoft Excel 2007 document206xmlXML document207xpiCompressed installation archive (XPI)208xzXZ compressed data209yamlYAML source210yaraYARA rule211zigZig source212zipZip archive data213zlibstreamzlib compressed data注意该清单是模型输出层所能输出的全部标签但并不是 Magika 工具最终可能返回的全部标签。诸如directory、empty、symlink、undefined等标签由工具层直接判定不经过深度学习模型详见下文“非模型路径”小节。二、从模型输出到最终标签一条标签的完整生命周期理解这份清单的意义关键在于弄清模型预测的标签如何变成你最终看到的label。核心实现在 magika.py 的_get_output_ct_label_from_dl_result()方法其处理顺序如下覆盖映射overwrite_map先将模型输出的标签经overwrite_map重写。standard_v3_0 的映射为{randomtxt: txt, randombytes: unknown, symlinktext: txt}——即模型在内部用三个额外标签细粒度建模随机文本、随机二进制与符号链接文本但对用户统一呈现为常规的txt/unknown。预测模式判定prediction_mode根据置信度分数与对应阈值决定是否采信模型结论。兜底降级若分数不足文本类文件回退为txt二进制类回退为unknown。2.1 三种预测模式PredictionMode预测模式通过 Python 接口的prediction_mode参数或 CLI 的--prediction-mode指定逻辑同样位于 magika.py模式判定条件行为BEST_GUESS无条件无论分数高低直接采信模型预测的标签HIGH_CONFIDENCE默认score thresholds[ct]无专属阈值时用medium_confidence_threshold分数达标才采信否则降级为txt/unknownMEDIUM_CONFIDENCEscore medium_confidence_threshold使用统一宽松阈值0.5判定分数不达标则降级2.2 standard_v3_0 的阈值配置来自 config.min.jsonmedium_confidence_threshold: 0.5绝大多数类型的默认高置信度阈值。thresholds: {handlebars: 0.9, latex: 0.95, markdown: 0.9, pascal: 0.95}这四类文本格式与普通代码文本容易混淆模型为它们单独设定了更苛刻的高置信度门槛降低误报率。min_file_size_for_dl: 8小于等于 8 字节的文件或去除空白后有效内容不足 8 字节不进入深度学习推理。padding_token: 256、block_size: 4096特征提取相关的填充标记与单次读取上限。beg_size: 1024、mid_size: 0、end_size: 1024v3.0 仅取文件开头 1024 字节与结尾 1024 字节作为模型输入不再取中段use_inputs_at_offsets: false表示不额外采样 0x8000/0x8800/0x9000/0x9800 偏移v2.1 同样为 false这些偏移位点是为 ISO/UDF 类文件预留的特征。特征提取的完整实现开头 lstrip、结尾 rstrip、padding 补齐等细节可参见 magika.py 的_extract_features_from_seekable()其要点是只读取文件开头与结尾各至多block_size字节避免将整个文件载入内存。2.3 非模型路径清单之外的标签以下标签不经过深度学习模型而是由工具层直接判定见 magika.pyempty0 字节文件score固定为 1.0。directory目录score固定为 1.0。symlink配合no_dereferenceTrueCLI 为-n/--no-dereference时不解析符号链接直接返回。undefineddl块中代表“未使用深度学习”例如过小文件、目录、空文件等场景此时dl.label为undefined而output.label为实际判定结果。txt/unknown文件过小 min_file_size_for_dl时直接尝试 UTF-8 解码可解码判为txt否则判为unknown见 magika.py。因此在实际使用中你会看到output.label的取值空间 上表 213 个标签 上述工具层标签。三、输出结构dl 与 output 的双层设计所有语言绑定Python、JS、Rust、Go 及 CLI返回统一的输出结构官方说明见 docs/magika_output.md。以识别一个 JS 文件为例$ magika tests_data/basic/javascript/code.js --json [ { path: tests_data/basic/javascript/code.js, result: { status: ok, value: { dl: { description: JavaScript source, extensions: [js, mjs, cjs], group: code, is_text: true, label: javascript, mime_type: application/javascript }, output: { description: JavaScript source, extensions: [js, mjs, cjs], group: code, is_text: true, label: javascript, mime_type: application/javascript }, score: 0.9710000157356262 } } } ]解读要点path本次预测对应的文件路径批量扫描多个文件时用于区分。result.statusok表示扫描成功文件不存在、权限不足等场景返回错误状态。score模型预测的置信度01。dl块深度学习模型的原始预测信息label对应上表 213 个标签之一未走模型时为undefined。output块Magika 工具层综合模型预测、置信度、预测模式后给出的最终结果是普通业务代码应直接使用的字段。is_text是否为文本类型决定低置信度时回退到txt还是unknown。官方建议“大多数客户端只消费output.label”原始dl预测主要用于调试。关于为何以label而非description/mime_type作为集成锚点可参见 docs/faq.md。四、在代码中查询与使用这套标签4.1 查询当前模型支持的完整标签列表Python 端提供了直接接口magika.pyfrom pathlib import Path from magika import Magika m Magika(model_dirPath(assets/models/standard_v3_0)) supported m.get_supported_content_types() print(len(supported)) # 213 print(supported) # 与 README 清单一一对应的 ContentTypeLabel 列表ContentTypeLabel枚举的完整定义见 python/src/magika/types/content_type_label.py其中包含全部候选标签模型实际只支持子集以上表为准。4.2 标签对应的元数据每个标签对应的描述、MIME 类型、扩展名、分组与is_text信息记录在知识库 python/src/magika/config/content_types_kb.min.json 中仓库根目录另有可读版本 assets/content_types_kb.min.json。模型推理时由_load_content_types_kb()载入magika.py文本类型默认 MIME 为text/plain二进制默认application/octet-stream未提供 description 时回退为标签名本身。4.3 实战解析一次识别结果from pathlib import Path from magika import Magika m Magika() # 默认加载 standard_v3_0 res m.identify_path(Path(tests_data/basic/javascript/code.js)) print(res.prediction.dl.label) # javascript模型原始输出 print(res.prediction.output.label) # javascript最终输出业务应使用 print(res.prediction.score) # 置信度如 0.971...上述测试样本位于 tests_data/basic/javascript/code.jsPython 端测试用例见 python/tests/test_magika_python_module.py。五、与其它内置模型的差异速览仓库内还提供了多个模型可通过model_dir参数切换模型输入特征标签数特性standard_v3_0默认beg 1024 end 1024213阈值含handlebars/markdown专属配置overwrite_map映射randombytes/randomtxt/symlinktextstandard_v2_1beg 2048 end 2048213不含randombytes/randomtxt/symlinktext标签overwrite_map为空fast_v2_1beg 512 end 512213输入更少速度优先见 assets/models/fast_v2_1/config.min.json三者的medium_confidence_threshold均为 0.5block_size均为 4096。模型的训练元信息如 100 个 epoch记录在 assets/models/standard_v3_0/metadata.json。六、小结standard_v3_0的 213 个输出标签构成了 Magika 文件类型识别能力的边界深度学习模型在这 213 个类别间做概率分布预测工具层再依据overwrite_map、专属阈值与预测模式做二次裁决最终通过dl/output双层结构同时暴露原始预测与可信结果。无论是直接阅读 assets/models/standard_v3_0/README.md 中的清单还是调用get_supported_content_types()在代码中动态获取你都可以精准掌握当前模型“能识别什么”从而在设计文件分类、内容安全扫描、上传校验等系统时做出正确的集成决策。【免费下载链接】magikaFast and accurate AI powered file content types detection项目地址: https://gitcode.com/GitHub_Trending/ma/magika创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考
返回列表