Interestana
Home/News/webAI Releases TwIL-LM Formal Logic Models for Local Autoformalization
MarkTechPost3 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

webAI Releases TwIL-LM Formal Logic Models for Local Autoformalization

webAI Releases TwIL-LM Formal Logic Models for Local Autoformalization

webAI has released TwIL-LM, a two-model family of formal-logic reasoners with parameter counts of 1.7 billion and 3 billion. The 3 billion parameter model, named TwIL-LM3, is a merged fine-tune of SmolLM3-3B, while the 1.7 billion parameter model is a PEFT LoRA adapter for SmolLM2-1.7B-Instruct. Both models are engineered to perform autoformalization, which involves translating English statements into first-order logic and verifying if a given conclusion logically follows from a set of premises. A key feature of the TwIL-LM family is its capability to run locally on user hardware. Quantized builds are available, with a 1.06 GB version for the 1.7B model and a 1.78 GiB Q4_K_M GGUF version for the 3B model. webAI's announcement highlights that TwIL-LM3 outperformed gpt-oss-120b on four out of five formal-reasoning benchmarks. However, the current release is restricted to non-commercial use under the webAI Non-Commercial License ver. 1.0, with revenue-generating deployments requiring a separate agreement with webAI. The models are suitable for companies of any size. The 3B Q4_K_M GGUF model requires 1.78 GiB of storage and can operate on a CPU or with as little as 4 GB of VRAM, while the 1.7B Q4_K_M model requires 1.06 GB. Industries that can benefit from this technology include compliance and RegTech, financial services, healthcare and pharmaceuticals, legal and contract operations, and formal-methods research. webAI positions the local execution capability as crucial for environments where sensitive data must remain on the device. Potential applications include first-order logic translation, entailment classification over premise sets, converting natural language into structured queries, assisting in drafting and critiquing Lean formalizations, and serving as a verifier layer to check the outputs of larger AI models. The development of TwIL-LM3 involved a four-stage process built upon the base model. This included supervised fine-tuning on a synthetic formal-logic corpus using LoRA, followed by checkpoint fusion which averaged intermediate SFT checkpoints in parameter space. A WiSE-FT interpolation step was then applied, interpolating back towards the pretrained base model with a lambda value of 0.25. The final stage involved MGPO, an entropy-weighted GRPO stage executed against a programmatic verifier. The published checkpoint corresponds to step 2071 of this process. The lambda value of 0.25 is significant, as it means only a quarter of the fine-tuned delta was retained. A parallel development path that omitted this interpolation achieved a higher in-domain score, reaching a macro gate of 0.515, but at the cost of approximately twelve points in held-out capability. webAI has not publicly released details on this trade-off.

Original source — read the full reporting at the publisher:

Read on MarkTechPost

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next