Add PATCH contact for missing peer IDs, contacts banner and edit UI,
enrich /demo with pairing steps, and update tester/signup guidance.
Co-authored-by: okuma <o0kuma@users.noreply.github.com>
Show/copy numeric user ID, require peer ID for contacts, start chat in one tap,
clarify empty states, and route L0 twin drafts into the human composer.
Co-authored-by: okuma <o0kuma@users.noreply.github.com>
Record invite/metrics/dashboard/draft(no_key) checks, pg_dump restore
rehearsal, and tester-facing guide for msn.iykyka.com.
Co-authored-by: okuma <o0kuma@users.noreply.github.com>
Attach web to nginx-proxy_default; record live cutover (NPM host #7,
old MSN stopped) and mark N2-B8–B12 done.
Co-authored-by: okuma <o0kuma@users.noreply.github.com>
msn.iykyka.com still serves legacy Express; agent lacks SSH/Portainer
credentials so cutover is documented and marked blocked pending access.
Co-authored-by: okuma <o0kuma@users.noreply.github.com>
Add Postgres + internal AI + core-backend + Flutter web (nginx API/WS
proxy). Document Portainer usage and env template; mark N2-B1–B7 done.
Co-authored-by: okuma <o0kuma@users.noreply.github.com>
Record local smoke results after restarting core-backend on current main.
Document msn.iykyka.com CORS/API change list for N1-6; next is N2-B Docker.
Co-authored-by: okuma <o0kuma@users.noreply.github.com>
Master: replace old MSN, AI internal-only, Web-first client, secrets in
Portainer/env only, ALLOW_DEMO_INVITE on for testers. N2-A complete.
Co-authored-by: okuma <o0kuma@users.noreply.github.com>
Consolidate Claude + Cursor DONE state and the next execution track
(smoke → Docker msn.iykyka.com → stabilize → FCM/Android QA → human PoC).
Sync stale PLANNING Q1–Q7/stack status and point AGENTS/README/CLAUDE/roadmap
at docs/deploy-checklist.md.
Co-authored-by: okuma <o0kuma@users.noreply.github.com>
- Resolve visual-direction conflicts (app_theme, main.dart, signup/onboarding
screens, index.html) in favor of the soft-gradient + glassmorphism design;
drop the competing Twin Shadow palette and TwinTokens usage.
- Port non-visual additions from the parallel branch: CORS middleware,
FlutterError/PlatformDispatcher crash handlers + boot timeout/fallback
screen in main.dart, and the shared demo-invite feature
(core-backend/demo.go, demo_test.go, mobile/lib/config.dart), reskinning
the demo panel to match the glass UI.
- Rename the shared demo invite code DEMO-BUNSIN -> DEMO-YKAVU on both
client and server so the tester-facing feature keeps working.
- Purge remaining "분신"/"bunsin" identifiers app-wide: Dart package name
(bunsin_mobile -> ykavu_mobile), Android applicationId/namespace
(com.bunsin.bunsin_mobile -> com.ykavu.ykavu_mobile, incl. Kotlin source
dir move), on-device DB filename, keystore alias/docs, PoC draft-generator
system prompt, and the static web boot placeholder div.
명칭 변경:
- decision-log.md Q6 확정: "분신"(가칭) → "와카뷰 (Ykavu)" — 焚身(분신자살)
동음이의 리스크도 있었고 메신저 브랜드로 부르기 무거웠음. Master가 최종
선택한 이름으로 변경(2026-07-31), 파생 문서(AGENTS.md/CLAUDE.md/PRD 등)·
Flutter 앱 텍스트·web manifest/index.html·AndroidManifest 라벨까지 전부 반영
- ai-service: 본인확인 고정 문구(identity.py)·시스템 프롬프트(generation.py)도
갱신 — 새 이름 기준으로 "본인이야 와카뷰야?" 류 질문을 감지하도록 정규식도
같이 손봄(단순 문자열 치환만으론 어미 형태가 안 맞아서 테스트 추가/조정)
- poc/tone-corpus/의 실험용 프롬프트는 의도적으로 그대로 둠(ai-service README에
이미 명시된 대로 프로덕션과 분리된 실험 도구)
UI/UX:
- mobile/lib/theme/app_theme.dart: 시드 컬러를 인디고/라벤더로 변경, 화면 전체에
깔리는 소프트 그라디언트(라이트: 라벤더→스카이→핑크 파스텔, 다크: 딥 인디고→
네이비→플럼) + 카드/인풋/칩을 반투명 "글래스" 서피스로 전환
- mobile/lib/widgets/gradient_backdrop.dart(신규): MaterialApp.builder에 연결해
모든 화면에 자동으로 그라디언트 배경 적용
- mobile/lib/screens/splash_screen.dart(신규) + main.dart 재구성: 기존엔
session.restore()를 기다리는 동안 아무 것도 안 그려서 흰 화면/텍스트만 뜨는
구간이 있었음 — runApp을 먼저 하고 restore 동안 브랜드 스플래시가 뜨도록 변경
- chat_screen·autonomy_settings_screen·data_flow_screen의 커스텀 패널들도
글래스 스타일로 맞춤
BunsinApp -> YkavuApp (mobile/lib/main.dart, test/widget_test.dart 동기화)
Flutter SDK가 없는 환경이라 flutter analyze/run으로 직접 컴파일 검증은 못했음 —
중괄호/괄호 균형과 기존에 검증된 API 패턴 위주로 신중하게 작성함. go test·
pytest는 전부 통과.
- Flutter: drift + SQLCipher local store for tone samples/KV with
secure-storage passphrase; migrate legacy SharedPreferences
- Core: FCM notifyUser on escalation, admin push-test, push metrics,
DELETE session for multi-device logout
- Docs/roadmap B checkboxes updated; Linux SQLCipher apt notes
Co-authored-by: okuma <o0kuma@users.noreply.github.com>
- Record draft/escalate latency and error rates on /admin/metrics
- Add minimal /admin/dashboard and full message JSON on send
- Fix identity answers in ai-service with stable copy + tests
- Flutter DataFlowScreen + SessionsScreen; device-token API stub
- Update roadmap B checkboxes after A3 API E2E
Co-authored-by: okuma <o0kuma@users.noreply.github.com>
Add conversation/contact/history/escalation-log endpoints, contact-scoped
whitelist matching, bearer sessions, ADMIN_API_TOKEN guards, and
`go run . migrate`. Update Flutter client to persist/send session tokens
and mark A1/A2 complete in the prioritized checklist.
Co-authored-by: okuma <o0kuma@users.noreply.github.com>
Scaffold `mobile/` against core-backend APIs (signup, chat, draft,
veto, retract, autonomy/whitelist). Update roadmap/README/AGENTS so
participant PoC #1/#3 stays the final Phase 1 step after app build.
Co-authored-by: okuma <o0kuma@users.noreply.github.com>
화이트리스트 CRUD:
- POST/GET /users/:id/whitelist-rules, DELETE /users/:id/whitelist-rules/:ruleId
- Flutter의 자율성 설정 화면(roadmap.md 2.3)이 바로 붙여 쓸 수 있게
준비. contact_id는 저장되지만 매칭 로직(whitelistMatches)은 아직
전역 키워드만 봄 -- 대화방-연락처 연결 모델링 필요(기존에 문서화된
한계, 그대로 유지)
되돌리기(one-tap undo, AGENTS.md 안전 불변식):
- Message.Retracted 필드 추가
- POST /messages/:id/retract -- 트윈이 자동발송한 메시지만 대상(사람이
쓴 메시지는 400), 이미 되돌린 건 409
- 성공 시 같은 대화방 WebSocket에 {"type":"retraction", "id":...}
브로드캐스트. 일반 메시지 브로드캐스트도 {"type":"message"}를 붙여서
클라이언트가 두 이벤트를 구분할 수 있게 함
되돌리기 버튼/사후알림 UI 자체는 여전히 Flutter 쪽 몫으로 남아있음
발견한 문제(2.6 진행 중): 거부권(peer veto) 안전 불변식이 코드에 전혀
구현되어 있지 않았음. Contact.TwinDisabledByPeer는 스키마에만 있고
어디서도 읽거나 쓰지 않았고, tech-design.md §4는 "대화방 단위 플래그"
라는데 실제로는 상대(Contact) 단위로 모델링돼 있어 설계 문서와도
불일치. Conversation.TwinDisabledByPeer로 옮기고 POST
/conversations/:id/veto 추가, 메시지 발송 시 거부권 -> 에스컬레이션
-> 자율성 레벨 순으로 체크(앞 단계가 뒤 단계를 항상 이김)하도록 수정.
2.6 본작업:
- 초대 기반 가입: 기존엔 아무 문자열이나 처음 쓰면 통과돼서 실제로는
초대 기반이 아니었음. InviteCode 테이블 + POST /invites(발급)
추가하고 /auth/signup이 미리 발급된 미사용 코드인지 검증하도록 변경
(모르는 코드 400, 이미 쓴 코드 409). 계정 삭제 시 코드는 "사용됨"
상태를 유지한 채 유저 참조만 지움
- GET /admin/metrics 추가 -- 메시지 수(휴먼/트윈), 에스컬레이션
사유별 집계, 거부권 발동률(vision.md 거부율 지표의 1차 근사),
초대 코드 발급/사용 수. 생성 지연시간·오류율은 별도 계측 계층이
없어 넣지 않고 문서에 명시
실제 베타 오픈 시점 자체는 roadmap.md §3(PoC 결과 필요)이 끝나야
정할 수 있어서 여전히 보류 -- 이번엔 서버 인프라만 준비함
ai-service:
- escalation_filter/retrieve_style/generation/main(FastAPI)의 ad-hoc
TestClient 검증을 ai-service/tests/ 정식 pytest 스위트로 승격(34개).
Gemini 호출은 mock, retrieve_style은 overlap이 recency를 항상
이긴다는 것(과거 버그 재발 방지)까지 포함
- on_event(deprecated) -> lifespan 컨텍스트 매니저로 교체
core-backend: 자율성 플로우(L0->L1->L2) 통합 테스트를 쓰려면 실제 분기
로직이 있어야 해서, roadmap.md §2.2에서 미결로 남아있던 자율성 엔진
오케스트레이션 최소 버전을 이번에 구현:
- PATCH /users/:id/twin-settings -- 자율성 레벨 변경 (기본값 L0)
- 트윈 발송 시 에스컬레이션 통과 후 레벨 확인: L0는 항상 차단, L1은
approved:true 필요, L2는 화이트리스트 매칭 시 즉시 자동발송·매칭
없으면 L1과 동일하게 승인 필요. 에스컬레이션은 레벨/화이트리스트/
승인 여부와 무관하게 항상 우선(테스트로 확인)
- 화이트리스트 매칭은 v1 최소 구현(전역 키워드 매칭, 상대별 예외는
아직) -- 대화방↔연락처 연결이 모델링되지 않아 보류, README에 명시
온보딩·채팅·설정 수동 QA는 Flutter 클라이언트가 없어 이 환경에서는
보류, roadmap.md에 근거 남김
발견한 문제: POST /conversations/:id/messages가 sender_mode=twin을
검증 없이 그대로 저장·브로드캐스트하고 있었음 -- 에스컬레이션 게이트는
초안 생성(/draft) 경로에만 있었고 실제 발송 경로엔 없어서, 클라이언트가
/draft를 거치지 않고 바로 twin 메시지를 보내면 안전선을 완전히
우회할 수 있었다.
- ai-service: /draft와 별개인 POST /escalate/check 하드게이트 엔드포인트 추가
- core-backend: AIServiceClient.checkEscalation 추가, 메시지 저장 직전에
twin 발송이면 무조건 호출하도록 해서 발송이 실제로 일어나는 단
하나의 지점에서 막음. AI 서비스 응답 불가 시 fail-safe로 발송 차단.
에스컬레이션되면 저장/브로드캐스트 없이 escalation_logs에만 기록.
사람이 직접 보내는 메시지는 게이트 대상 아님
- core-backend: DELETE /users/:id 추가 -- 유저가 걸린 모든 행(트윈 설정·
화이트리스트·연락처·대화참여·메시지·에스컬레이션로그·유저 본인)을
트랜잭션으로 삭제 (tech-design.md §5 "사용자가 언제든 초기화 가능")
- 온디바이스 암호화·데이터 흐름 대시보드는 Flutter 클라이언트 책임이라
이 환경에서는 보류, roadmap.md에 근거 남김
ai-service/ wraps generate_draft/escalation_filter/retrieve_style
behind a single POST /draft endpoint that the Go core will call
internally. poc/tone-corpus/ stays untouched for corpus experiments
and blind-eval; this is the promoted copy meant for the real service.
Verified with TestClient: style_examples path, history/retrieval
path (confirms the earlier scoring fix still ranks the on-topic
exemplar first), escalation short-circuit, and 422 validation when
zero or both of style_examples/history are given.
Still missing: the Go core's actual HTTP client calling this service.
2.3 (Flutter) needs a Flutter/Dart SDK this environment doesn't
have, so it can't be built or verified here the way core-backend
and the AI pipeline can. Swapping it with 2.2 (Python AI service)
keeps work unblocked instead of writing unverifiable Flutter code;
2.3 moves to wherever the user has the Flutter SDK installed.
Ports the Python prototype (backend/) to the actual chosen stack --
Gin + gorilla/websocket + GORM, same DB schema (models.go mirrors
backend/app/models.py), same endpoints (signup, message send,
WebSocket relay). backend/ stays as a reference prototype, not
removed.
Verified with go test: signup, duplicate-invite-code rejection (409),
404 on an unknown conversation, and WebSocket broadcast delivery all
pass -- the same cases the Python version was checked against.
Push notifications, AI service integration, and multi-device sync
are not in this commit -- see core-backend/README.md.
The build order already ended with the PoC-gated section, but it
wasn't stated as a hard rule. Now roadmap.md and AGENTS.md both say
not to touch Phase 1 §3 until items 1-5 of the build order are fully
finished, even if PoC data happens to land earlier -- no jumping the
queue to fill in a value early.
Flutter weakens the old "single platform = faster dev" argument for
skipping iOS, so tech-design.md §6/§8 now state the reasons that
still hold: v2's OS-layer notification access is Android-only by
Apple's own policy regardless of framework, and the dev environment
is Windows, so an iOS build isn't even possible right now. Flutter
still means no UI rewrite once a Mac is available later.
Reconsidered pure Go (would mean reimplementing and re-verifying the
already-tested AI pipeline) and pure Python (leaves perceived
performance/concurrency headroom on the table for a preemptive
scale bet). Landed on: Go handles auth/messaging/DB, Python keeps
owning generate_draft/escalation_filter/retrieve_style behind an
internal API. backend/ (Python) is now labeled a verified prototype
for the Go rewrite to match, not the final service.
Native wasn't actually required for the v2 OS-layer notification
listener -- Flutter reaches it via platform channels like any other
native Android API, same pattern many production apps already use.
Flutter's faster iteration on the chat UI (already validated in the
click prototype) matters more for v1 than starting native. Updates
tech-design.md §8, roadmap.md's checklist, and AGENTS.md accordingly.
FastAPI app with invite-code signup, message send/relay over
WebSocket, and the DB schema from roadmap.md Phase 1 §2.1 (users,
contacts, conversations, messages, twin_settings, whitelist_rules,
escalation_logs). Defaults to SQLite for local dev, PostgreSQL in
prod per tech-design.md §8.
Verified end-to-end with TestClient: signup, duplicate-invite-code
rejection (409), message persistence, 404 on an unknown conversation,
and WebSocket broadcast delivery all behave as expected.
Push notifications and the AI pipeline integration (item 3) are not
in this commit -- see backend/README.md and roadmap.md's checklist.
tech-design.md §8 settles the stack (Android/Kotlin, Python/FastAPI,
PostgreSQL, WebSocket relay, Room+SQLCipher) so it stops blocking
item 1 of the build order. roadmap.md's Phase 1 breakdown is now
checkboxes instead of prose, and AGENTS.md adds the rule to check/
update that checklist before and after any Phase 1 app-build task,
rather than tracking progress ad hoc.
PoC execution is on hold for now, so this splits Phase 1 into work
that's independent of PoC results (backend infra, client shell, AI
pipeline productionization) vs. values that genuinely need PoC data
(autonomy defaults, whitelist topics, trust UX copy) -- so
infrastructure work isn't blocked while PoC recruiting catches up.
blind_eval.py runs PoC #1's blind-eval methodology over N held-out
corpus dialogues automatically (generate_draft.py refactored to
expose draft_reply() so both share the same drafting logic).
retrieve_style.py implements the keyword/recency search from
tech-design.md §2-1 and wires into generate_draft.py as --history,
replacing hand-curated --style files.
Bash access was intermittently restricted for part of this session
(auto-mode safety classifier), so these were initially written and
committed-pending without live execution. Now verified for real:
generate_draft.py's existing behavior still holds after the
draft_reply() refactor, blind_eval.py runs cleanly against val.jsonl,
and retrieve_style.py's original weighted-sum scoring had a real bug
-- recency drowned out keyword overlap for short Korean messages
(particle attachment means "핀란드" and "핀란드는" don't share a
token), so it was effectively returning the most recent messages
regardless of topic. Fixed by ranking on (overlap, recency) instead
of a weighted sum, confirmed the Finland-related exemplar now ranks
first for a matching query.
escalation_filter.py implements tech-design.md §3's first step as an
actual hard gate, not just a system-prompt instruction: money,
appointment-confirmation, and emotional content stop generate_draft.py
before it ever calls Gemini. Self-test 10/10; measured a 0.93% trigger
rate against 82,305 real corpus utterances (mostly factual price
mentions, not personal money requests -- noted as an upper bound, not
a real-usage estimate).
v1 doesn't train a custom model: it retrieves the closest-matching
past messages from the person's own history and feeds them as
few-shot exemplars to the same prompt contract generate_draft.py
already implements, via the hosted Gemini call. Narrows the AI-Hub
base corpus's role to evaluation and future on-device distillation,
since a hosted LLM already covers general Korean fluency.
build_unlabeled_corpus.py pulls the 34,030 raw-source dialogues that
never got speech_act/slot labels (found via the earlier QA
cross-check) into their own unlabeled.jsonl, forward-filling the
per-dialogue metadata that AI-Hub's CSV export only writes on each
dialogue's first row. Verified against the actual data: 176,605
raw - 142,575 labeled = 34,030, matches exactly.
generate_draft.py takes a few style-exemplar messages plus recent
conversation context and drafts a reply via LLM call (the server
fallback path from tech-design.md §2), with escalation baked into
the system prompt for money/appointment/emotional content. Verified
prompt construction against a real corpus dialogue and hand-compared
a generated draft to the withheld real reply (README "샘플 검증") --
no API key in this session, so the live call itself is untested.
Consolidates the AI-Hub "한국어 SNS 멀티턴 대화" TL/VL zip parts into
clean train/val JSONL (142,575 dialogues), with a QA cross-check
against the raw TS/VS source. Script and docs only -- the dataset
itself stays out of git per .gitignore, both for size and because
AI-Hub's terms restrict redistribution.
Consolidates what's ready (doc set, PoC plans/materials, prototype)
vs. what still needs a human to execute (recruiting, interviews,
Go/No-Go calls), plus a Q1-Q7 confirm/revisit checklist for the
actual planning meeting.
poc-materials.md only tests how the recipient reacts to the twin; Q3
asks whether the person delegating to it actually wants to. Adds a
separate screening + interview script for that side, with guidance
on sequencing it before the peer role-play so the two perspectives'
gap becomes a signal in itself.
Recruiting message and data-consent blurb for PoC #1, a role-play
script for the one scenario the click prototype doesn't cover
(emotional escalation), and a shared post-session interview guide —
so participant recruitment can start without drafting these from
scratch.
Built an interactive click-through prototype covering the read-receipt,
identity-confirmation/veto, and escalation scenarios plus the autonomy
settings screen. Link it into PLANNING.md's checklist and poc-plan.md
so PoC #3 role-play can reuse it as stimulus material instead of
building scenario scripts from scratch.
Turns the two highest-priority open risks (on-device tone realism,
impersonation/trust acceptance) into runnable protocols: sample
collection, blind evaluation, role-play scripts, and pass/fail
thresholds, so results can update decision-log.md and risk-log.md.